Discussion about local AI usually begins with the model: its size, speed, memory use, and benchmark score. For a narrow product, that is the least stable part of the design. Models change. The task should survive them.
The durable asset is a controlled teaching loop. It defines what the system must learn, how success is measured, and what happens when it fails. A model enters that loop as a candidate implementation.
A larger model prepares the lesson
A capable model can be used during development as a teacher. It reads the source material, identifies the relevant concepts, and produces examples in the small language expected by the product. It can generate variations, counterexamples, and focused corrections. None of this requires the teacher to be present after deployment.
The student receives a much smaller job. It maps new material into the bounded representation it has learned. Deterministic software takes over from there. Structure, identifiers, layout, packaging, and syntactic validity do not need another act of interpretation.
The teacher can build the harness too
The same teacher can draft the machinery used to evaluate the student: expected outputs, malformed cases, edge cases, scoring rules, and tests aimed at known weaknesses. This is the useful extension. AI contributes to the classroom as well as the lesson.
Origin does not grant authority. Once reviewed, a generated test becomes a fixed artifact. It must reject a known bad output. A structural validator must produce the same verdict every time. Evaluation cases used for final measurement remain outside the training material.
This gives the loop a simple rhythm. Train the student. Run the harness. Inspect the failure. Ask the teacher for material aimed at that failure. Admit the useful material into the corpus. Train again. Each cycle has a reason and leaves a trace.
Failure improves the system
A failed generation is useful when it changes the next lesson. Missing concepts create new examples. Broken formatting creates stricter cases. Repetition creates counterexamples. The corpus grows from observed weakness instead of volume for its own sake.
Progress becomes visible. The system either passes more held-back cases or it does not. Execution time and artifact validity can be measured. A model that stops improving can be replaced without discarding the grammar, corpus, validators, or history of failures.
The human owns the grammar
The teacher can propose the harness, but it cannot certify its own interpretation. A fluent mistake can be copied into the examples, learned by the student, and rewarded by an evaluator built from the same assumption.
The human boundary is narrow and consequential: choose the legitimate sources, define the target grammar, decide which failures matter, review changes to the corpus, and keep part of the evaluation independent. Human attention governs meaning rather than supervising every generated token.
That boundary is the architecture. It allows AI to generate training material and parts of its own harness without allowing generation to redefine success.
A small local model is therefore not a diminished version of a frontier model. It is the replaceable worker inside a teaching system. The product is the loop that can teach another worker tomorrow and still know whether the work is acceptable.
Research basis
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, 2015. Distilling the Knowledge in a Neural Network. A foundational description of transferring learned behaviour from a larger or combined teacher into a smaller model suited to deployment.
Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo, 2015. The Oracle Problem in Software Testing: A Survey. A survey of the central testing problem of determining whether an observed result is correct, including specifications, contracts, models, and other sources of test-oracle information.
Yizhong Wang and colleagues, 2023. Self-Instruct: Aligning Language Models with Self-Generated Instructions. A demonstration that model-generated instructions and examples can form useful training material when they are generated, filtered, and used through a defined process.
Nadia Alshahwan and colleagues, 2024. Automated Unit Test Improvement using Large Language Models at Meta. An industrial study of LLM-generated test improvements filtered through compilation, reliability, and coverage checks before being proposed for acceptance.
Koki Wataoka, Tsubasa Takahashi, and Ryokan Ri, 2024. Self-Preference Bias in LLM-as-a-Judge. An examination of systematic bias that can appear when language models evaluate generated responses, supporting caution about circular, model-only acceptance.