Booking · Q4 2026 · 2 slots open20+ yrs shipping systemsIdea → prototype → production
All guides
September 29, 20269 min readWhere AI Is Worth It

The Loop That Improves Itself: LoRA Fine-Tuning on Your Agent's Real Failures

The loop that improves itself: turn an agent's flagged failures into training data for a small, LoRA fine-tuned model, at NVIDIA's own numbers.

Collect every response your grading signal flags as wrong instead of discarding it, then split those failures into named categories before fine-tuning anything. NVIDIA replaced a 70B router with an 8B model that way, matching its accuracy at 10x smaller and 70% lower latency.

You need a grading signal somewhere in your system before this loop matters: a human reviewer, a rubric, or an LLM-as-judge grading a sample of traffic. Grading an Agent is where that comes from. If you're not sure you have an agent worth grading in the first place, that question comes before this loop.

What Is the Loop That Improves Itself?

I use this as shorthand for a specific build:

  1. Collect the failures a system already produces.
  2. Sort them by cause.
  3. Fine-tune a small model on the worst one.

You'll sometimes see this pattern called AI distillation. NVIDIA's own paper never uses that word for the mechanism behind it. What it describes is LoRA fine-tuning: training a small set of extra weights on a curated set of real examples, instead of retraining the whole model or copying a bigger model's output probabilities.

Which Failures Should You Keep?

A negative sample is a real interaction your grading signal flagged as wrong, kept instead of thrown away. A response a real user got, that a human or a judge marked as a miss, never a hypothetical you invented or a test case you wrote.

NVIDIA's own internal assistant, NVInfo AI, is a mixture-of-experts system serving more than 30,000 employees. Over a 3-month window after it shipped, NVIDIA's team kept every negative signal from real usage instead of discarding it. That produced 495 negative samples: real interactions from actual usage, kept on purpose.

495 sounds small against more than 30,000 people using the tool. You don't need a huge dataset for this. You need real mistakes, kept on purpose, over a long enough window to see the same one happen more than once.

Why Sort Failures Into Categories First?

NVIDIA didn't retrain on 495 things going wrong as one bucket. Its team split the 495 negative samples by cause. Twenty-six were routing errors, the wrong specialist model picked for a query. About 16 were query-rephrasal errors, where reformulating the question lost its original meaning.

Category Share of the 495 What went wrong
Routing errors 5.25% (26 samples) The wrong specialist model got the query
Query rephrasal errors 3.2% (about 16 samples) Reformulating the question lost its meaning
Everything else (retrieval, ranking, generation) Over 90% Untouched by this particular fix

Routing and rephrasal errors combined made up less than 10% of the 495 negative samples. NVIDIA fixed the two categories with a name and a number attached, not the other 90%.

Why Fine-Tune Small Instead of Scaling Up?

The router's training set was not the raw 495 failures. NVIDIA's team built a separate, curated set from them: the 26 confirmed routing errors, plus correct examples and expert-corrected completions, for 761 data points, 685 left after removing duplicates.

They fine-tuned an 8B model on that set with LoRA, a method that trains a small set of extra weights instead of retraining the entire model. The result: the 8B model matched the 70B router's accuracy exactly, both scoring 96%, at 10x smaller and 70% lower latency.

NVIDIA's team ran the same method on the rephrasal problem: 10 error patterns pulled from 250 hand-reviewed samples seeded roughly 5,000 synthetic examples, generated with a larger model. They got a 3.7% accuracy gain and a 40% latency drop from that rephrasal run.

How Do You Catch a Misquoted Number?

NVIDIA's router work gets repeated in secondary write-ups as a flat 98% cost savings. That number is real. It is also not about the router above at all.

NVIDIA runs a second, related open-source project called the Data Flywheel Blueprint, described in its own GitHub README [3]. That README describes a different case entirely: an internal HR chatbot, where a fine-tuned 1B model reached about 98% of a 70B model's accuracy on tool-calling.

A separate line in the same README says the Blueprint can cut inference cost by up to 98.6% in the best case. Both README figures, the ~98% accuracy on tool-calling and the up-to-98.6% cost reduction, describe the HR chatbot case. Neither is about the 8B router discussed above.

The arXiv paper's own Table III states the actual number for that router: the 8B and 70B models scored identically, 96% each.

Before repeating a number that sounds too clean, find where it was first published and confirm exactly which claim it is attached to. The arXiv paper covers the NVInfo AI router. The README covers a separate build, the HR chatbot and the Blueprint.

Does the Loop Keep Paying Off?

NVIDIA ran this pass once and got a better router. Meta's Segment Anything data engine shows the same loop compounding, each turn cheaper than the last.

In the first stage, annotators corrected an early version of the model's mask predictions by hand. As the model improved inside that stage, annotation time per mask fell from 34 seconds to 14 seconds, producing 4.3 million masks.

A middle, semi-automatic stage let the model propose confident masks on its own so annotators could spend their time on what it still missed. By the final, fully automatic stage, the model needed no correction at all, and the dataset reached 1.1 billion masks across 11 million images.

Meta's annotators got roughly 2.4x faster inside the first stage alone, because the model handled more of the confident masks on its own. That freed them to spend their time on what it still missed, which produced a bigger, cleaner training set for the next version of the model.

What If the Model Grades Its Own Work?

Google DeepMind ran a different experiment, not NVIDIA's and not Meta's: what happens when a model corrects its own answer with no outside signal at all [4]. On GSM8K math problems, GPT-3.5's self-correction barely moved accuracy: 75.9% to 75.1% to 74.7%.

On CommonSenseQA, the same process collapsed accuracy from 75.8% to 38.1%, recovering only to 41.8% on a second attempt. A 34-point loss, self-inflicted.

Every signal feeding the loop came from outside the model being fixed. NVIDIA's team read real thumbs-down feedback. Meta's annotators corrected real masks by hand.

Run This Checklist This Week

  1. Start keeping every response your grading signal marks wrong, instead of discarding it.
  2. Once you have a real set, split it into at least two named failure categories, not one aggregate rate.
  3. Before assuming you need a bigger model, check whether fine-tuning a small model on your worst category alone would close the gap.
  4. Before repeating a published number from a case like this one, find its primary source and confirm exactly what it measures.

Quick Recap

  • Keep every response your grading signal flags, instead of discarding it. A negative sample is a real response marked as wrong. NVIDIA kept 495 of them over 3 months.
  • Sort before you fine-tune. NVIDIA's team fixed the two named categories: routing errors (5.25%) and rephrasal errors (3.2%). The other 90% of failures stayed untouched that run.
  • A small model can match a big one. NVIDIA's 8B router matched its 70B predecessor's accuracy exactly, at 10x smaller and 70% lower latency.
  • "Distillation" is the term secondary write-ups use. NVIDIA's own paper never uses it. What shipped was LoRA fine-tuning on a curated 685-example set.
  • A 98% figure and a 96% figure can both be real and still describe two different NVIDIA projects. Check the primary source before repeating either one.
  • You get compounding returns when the signal comes from outside the model being fixed. Meta's annotators got roughly 2.4x faster as the model improved. When Google DeepMind ran self-correction with no outside signal, accuracy fell instead.

Start Here

The intake at daisyguti.ai/work-with-me is about nine questions and takes a few minutes, held as a live conversation with an AI. It turns your answers into a brief with the specifics that matter:

  • What you're trying to solve
  • What you've already tried
  • What a working version looks like to you

I read that and reply with whether the fix is a bigger evaluation layer first, or a fine-tuning pass like the one above.

The AI buzzwords guide has plain definitions for the vocabulary this piece assumes, and choosing a retrieval cutoff is the same kind of named threshold, built for a different part of the pipeline.

Sources

  1. Shukla, Knowles, Madugula, Farris, Angilly, Pombo, Xu, An, Balasubramanian, Yu, Ren, Akkiraju (NVIDIA), "Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement," arXiv:2510.27051 - https://arxiv.org/abs/2510.27051 - verbatim: "NVInfo AI, NVIDIA's Mixture-of-Experts (MoE) Knowledge Assistant serving over 30,000 employees"; "we monitored feedback and collected 495 negative samples" over "a 3-month post-deployment period"; "routing errors (5.25%)" and "query rephrasal errors (3.2%)"; "we replaced a Llama 3.1 70B model with a fine-tuned 8B variant, achieving 96% accuracy, a 10x reduction in model size, and 70% latency improvement"; "fine-tuning yielded a 3.7% gain in accuracy and a 40% latency reduction"; the router's training set was "761 data points, consisting of 729 original samples and 32 additional corrections generated by the LLM-as-Judge," reduced to "685 unique samples" after deduplication; the rephrasal set used "10 key error patterns from 250 manually reviewed samples" to generate "approximately 5,000 rephrased queries" via Llama 3.1 405B; routing and rephrasal errors together "made up less than 10% of all system failures."
  2. Kirillov, Mintun, Ravi, Mao, Rolland, Gustafson, Xiao, Whitehead, Berg, Lo, Dollár, Girshick (Meta AI Research), "Segment Anything," arXiv:2304.02643 - https://arxiv.org/abs/2304.02643 - verbatim: "Average annotation time per mask decreased from 34 to 14 seconds" in the assisted-manual stage, producing "4.3M masks from 120k images"; in the semi-automatic stage the model "automatically detected confident masks" and annotators were "presented with images prefilled with these masks and asked them to annotate any additional unannotated objects," adding "an additional 5.9M masks in 180k images"; the fully automatic stage applied the model "across all 11M images" to produce "1.1B high-quality masks."
  3. NVIDIA-AI-Blueprints, "data-flywheel" (GitHub repository README) - https://github.com/NVIDIA-AI-Blueprints/data-flywheel - verbatim: on an internal HR chatbot case, "a fine-tuned llama-3.2-1b-instruct was able to achieve ~98% accuracy relative to the 70b model"; "using a Data Flywheel can reduce inference costs by up to 98.6%, while maintaining comparable accuracy."
  4. Huang, Chen, Mishra, Zheng, Yu, Song, Zhou, "Large Language Models Cannot Self-Correct Reasoning Yet," ICLR 2024, arXiv:2310.01798 - https://arxiv.org/abs/2310.01798 - "LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction"; Table 3 reports GPT-3.5 intrinsic self-correction on GSM8K at 75.9% to 75.1% to 74.7%, and on CommonSenseQA at 75.8% to 38.1% to 41.8%.

Start a project

Ready to scope it and ship it?

Have an AI workflow, reporting gap, or system you can't seem to get off the ground? Scope it with me. You'll get a real read on what to build first and what it should cost.

I review every submission and reply with a real read on fit.