Home
WorkBlogTalksPapersAbout
Nikita Kozodoi

Nikita Kozodoi, PhD

Senior Scientist on the AWS FDE team. I publish applied AI research and embed with engineering teams to build custom AI systems, from scoping and experimentation to fine-tuning and production deployment.

8+ years in applied AI/ML · Focused on LLMs and agentic systems

CV
29
Blog Posts
35
Public Talks
17
Papers
615
Paper Citations
397
GitHub Stars
18
Kaggle Medals

Latest work

Latest Blog

SWE-InfraBench + Kiro: What Happens When a Coding Agent Tackles IaC?

Builder Center

Language models solved only 34% of SWE-InfraBench's infrastructure-as-code tasks in one attempt. We ran Kiro over all 100 and found that a full agent loop with Sonnet 5 reaches 82%, with partially correct answers converging once the agent can re-run the tests itself.

View all 29 blog posts
Latest Talk

Intelligent Document Processing at Scale with the IDP Accelerator

GITEX AI · Berlin

Live demo of the IDP Accelerator, a scalable, serverless solution for automated document processing and information extraction on AWS. The asset combines generative AI and optical character recognition (OCR), using services such as Amazon Bedrock Data Automation and Amazon Bedrock foundation models to extract, classify, and process documents at scale.

View all 35 talks
Latest Paper

Are We Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

ICML 2026 Workshop on Weight-Space Symmetries

Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of domain experts affects the quality of the merged model. We fine-tune experts on five domains (Math, Code, Instruction Following, Multilingual, and Safety) across three model sizes (Qwen 3.5 0.8B, 2B, and 4B), saving checkpoints from 25% to 500% of the optimal training steps and evaluating five merging methods at each duration. Our findings reveal a striking method-dependent pattern: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum. We formalize this through bias-variance decomposition analysis, drawing a parallel to random forests where averaging benefits from high-variance individual learners. These results suggest that training duration and merging method should be chosen jointly rather than independently.

View all 17 papers

Selected work

All 8 case studies
AUMOVIO

Multi-agent defect detection for automotive software

Agentic code scanning system that found 30+ high-priority defects in a production automotive codebase.

Toyota Motor Europe

Recovering business logic from a mainframe

Agentic pipelines that recovered the business logic buried in 1.3 million lines of mainframe code, now moving to production rollout.

Lounge by Zalando

Localizing marketing content across 17 European markets

Test-time self-reflection for content localization across 17 European markets, tuned for quality, cost and latency.

Get in touch

Open to collaboration, speaking invitations, and conversations about applied AI.

ContactSee my work

Pages

  • Work
  • Blog
  • Talks
  • Papers
  • About

Social

  • LinkedIn
  • GitHub
  • Google Scholar
  • X / Twitter
  • Instagram

Contact

  • n.kozodoi@icloud.com
  • Buy me a coffee
  • Download CV
  • RSS feed
  • Berlin, Germany

© 2026 Nikita Kozodoi. All opinions are my own.