ENGINEERING LLM EVALUATION 2026-08-07

Beyond the vibes

The AI industry has a measurement problem. For natural language, "LLM-as-a-Judge" makes sense: a larger model scores a smaller model's output against a qualitative rubric. But code is not a poem. It compiles, it has defined edge cases, and its runtime complexity scales objectively. Asking an LLM "does this code look right?" is an engineering anti-pattern. Drawing on two decades of software engineering and formal automata theory, I built and open-sourced an alternative: Acceptance Test-Driven LLM Evaluation (ATD-Eval) - a deterministic 3-tier gate in Node.js that treats models the way compilers treat humans.

The problem

Vibes aren't verification

A judge LLM will hallucinate a passing grade if the variable names simply look convincing.

The gate

Three deterministic tiers

AST structure → sandboxed unit traps → empirical Big-O profiling. Any gate fails, the pipeline stops.

The result

40% → 100% on trap cases

Injecting algorithmic theory into context before generation beat every self-correction loop.

# The ATD-Eval pipeline, in one block [LLM output] → Tier 1: AST inspection # structure: DP loop or smuggled recursion? → Tier 2: vm unit execution # logic: survives the [1,3,4] greedy trap? → Tier 3: Big-O profiling # scaling: O(N·K) verified, O(2ⁿ) rejected [!] Any gate fails → pipeline terminates. No judge, no rubric, no vibes. _

01· The litmus test: memorization vs. reasoning

Most LLMs ace textbook algorithms because they retrieve them from training data, not because they reason through them. To test logical adherence, you must force the model out of its latent space. Classic algorithms curricula - Anany Levitin's Introduction to the Design and Analysis of Algorithms, or for readers from the Balkans, Dejan Živković's Uvod u algoritme i strukture podataka - highlight a perfect trap: the Change-Making (Coin Change) problem.

With canonical coins [1, 5, 10, 25], a Greedy algorithm works perfectly. Introduce a non-canonical set like [1, 3, 4] with target amount 6, and Greedy fails: it picks 4+1+1 (3 coins) instead of the optimal 3+3 (2 coins). Solving it correctly requires a bottom-up Dynamic Programming state table.

So if your prompt strictly specifies "write a bottom-up DP array; no recursion, no greedy logic" - how do you mathematically verify the LLM listened? Regex isn't powerful enough, and an LLM-as-a-Judge will often hallucinate a passing grade when the code merely looks convincing. That question is what the three gates answer.

02· The 3-tier deterministic gate

Tier 1 — Static AST inspection (structural adherence)

Regex parsing is notoriously brittle for code analysis. Instead, the TypeScript Compiler API parses the LLM's output into an Abstract Syntax Tree before execution. Walking the AST verifies constraints programmatically: Did it allocate an array for DP state? Did it use an iterative loop? Did it sneak in an unauthorized recursive call? Structural violations terminate the pipeline immediately - no CPU cycle is ever spent executing code that already broke the contract.

// Tier 1: walking the AST for structural violations function visit(node: ts.Node) { if (ts.isCallExpression(node) && node.expression.getText(sourceFile) === 'coinChange') { isRecursive = true; // caught unmemoized recursion } if (ts.isArrayLiteralExpression(node) || ts.isElementAccessExpression(node)) { usesArray = true; // confirmed DP state table } ts.forEachChild(node, visit); }

Tier 2 — Sandboxed unit execution (logical adherence)

If the structure is sound, the TypeScript is compiled in-memory and executed in an ephemeral context via Node's vm module with a hard CPU timeout - LLMs occasionally hallucinate infinite loops, and an unbounded main thread means a frozen runner. (Stated plainly: vm is a compartment, not a security boundary - Node's own docs say so. For genuinely hostile code you'd add process-level isolation; for an evaluation harness, a fresh context plus a hard timeout is the right tool.) Here the [1, 3, 4] trap instantly fails models that fell back on their pre-trained Greedy bias.

Tier 3 — Empirical Big-O profiling (complexity adherence)

Passing unit tests isn't enough: a poorly optimized recursive solution might pass Tier 2 and still explode at O(2ⁿ) in production. The pipeline load-tests the V8 engine at N=1,000, N=10,000, and N=50,000. Crucially, it runs warm-up iterations first - without warming the V8 JIT, your timing measures the compiler's optimization overhead, not the algorithm. Once warmed, if a 5x input increase causes an exponential execution spike, the solution is rejected for violating the O(N·K) constraint.

03· Benchmark: context beats correction

Because LLMs are stochastic, a single execution proves nothing. I ran the pipeline with Pass@k sampling (k variations at non-zero temperature) on Gemini 2.5 Pro using two prompt strategies: a baseline constraint ("Solve using DP. No Greedy.") and a theory-injected context - RAG-style pre-conditioning that explains why Greedy fails on [1,3,4] before generation.

Pipeline stage Baseline Theory-injected
Tier 1 · AST structure 70% 100%
Tier 2 · Unit accuracy 40% 100%
Tier 3 · Big-O profiling 30% 90%

Without theoretical grounding, the model's latent bias toward memorized Greedy solutions fought the prompt instructions - producing recursive solutions that failed the AST gate, or falling into the trap despite explicit instructions. Pre-conditioning the context altered the model's reasoning path at the source and eliminated the need for a self-correction loop entirely.

04· Why massive prompts fail

Scale the requirements up and another systemic issue appears. Dump a massive specification into a 1-million-token context window and the model will often hallucinate. Researchers at Stanford and UC Berkeley identified the "Lost in the Middle" phenomenon (Liu et al., TACL): Transformer attention has primacy and recency bias - it reliably remembers the beginning and end of a prompt, while facts buried in the middle are mathematically diluted into noise.

The antidote is the Single Responsibility Principle, applied to prompting: bound the search space by breaking work into atomic chunks. This is exactly how multi-agent frameworks like LangGraph and AutoGen operate under the hood - a Planner decomposes the spec, an Executor runs one task in total isolation with a dense, focused context, and failed tool executions route the stack trace back for a self-correction loop.

05· RAG vs. fine-tuning: the open-book exam

When systems demand absolute factual and logical accuracy, the standard debate arises: fine-tune, or use Retrieval-Augmented Generation? Think of it this way: fine-tuning is studying for a closed-book exam; RAG is an open-book exam. The widespread belief that fine-tuning teaches a model new facts is a misconception - it suffers catastrophic forgetting and leaves no audit trail when the model eventually hallucinates. Fine-tuning is for form and behavior (an XML schema, a brand voice).

RAG is for logic and facts. Storing verified algorithms and constraints in a vector DB and injecting them into context right before execution is exactly what the benchmark above measured - and it took Tier 2 accuracy from 40% to 100%. The engineering standard is clear: RAG for dynamic knowledge retrieval; fine-tuning only to teach the model how to perfectly format what it retrieved.

06· Three rules from the bench

  • 1> Grade code with computation, not conversation. If a claim about code can be checked by a compiler, a sandbox, or a profiler, an LLM's opinion of it is noise.
  • 2> Trap the memorization. Standard benchmarks measure retrieval. Non-canonical edge cases - like [1,3,4] - measure reasoning.
  • 3> Fix the context, not the model. Theory injected before generation outperformed every after-the-fact correction loop we ran.

The shift from probabilistic text generation to deterministic software engineering requires a new testing paradigm. ATD-Eval is a step in that direction - the full 3-tier engine is open source, and if you're building autonomous agents or evaluation pipelines, I'd love to hear how you're handling deterministic testing.