AI & DevelopmentDeveloper Tools

D-Engine: 14x Fewer Tokens, Zero Broken Commits

On 10 benchmark tasks, D-Engine matched a full agentic loop’s quality score while consuming 2,100 tokens instead of 93,000 and completing in 2.7 seconds instead of 38. No trade-off in output quality. Zero broken commits compared to one for the agentic approach. That result — published this week alongside the MIT-licensed code — makes a blunt argument: for well-scoped code edits, single-pass determinism beats open-ended agency.

D-Engine is a TypeScript harness built by corruchaga on GitHub that restricts what the LLM can do. The model receives a description of the desired change and target files, then returns only SEARCH/REPLACE blocks — no prose, no full-file rewrites. A local runtime applies those blocks in a shadow git worktree, runs tsc --noEmit, and merges to your main branch only if compilation passes. The LLM is invoked exactly once per task.

How the Compilation Gate Works

The shadow git worktree is the key safety mechanism. Every proposed change lands in an isolated branch that your working tree never sees until the gate clears. A malformed edit — wrong indentation, a missing import, a mismatched brace — fails the compile check and disappears. Your main branch stays clean regardless of what the model produces.

The patching process uses a four-tier cascade to apply the model’s SEARCH/REPLACE blocks:

  1. Exact string match
  2. Normalized newlines
  3. Trailing whitespace tolerance
  4. Fuzzy matching at a 0.85 similarity threshold

Three safety modes let you tune risk tolerance. Fast applies changes directly. Verify applies and checks before committing. Shadow — the recommended default — requires explicit user approval before anything merges. D-Engine’s own framing sums it up: “the AI thinks, the gate decides.”

The Benchmark Numbers

The 10-task benchmark used a frozen task set, the same model family, and one attempt per task for both approaches:

MetricD-EngineAgentic Loop
Quality (max 50)4848
Avg tokens~2,100~93,000
Avg time2.7s38s
Broken commits01

These numbers align with broader harness research. Vercel found that removing 80% of available tools from their agent raised success rates from 80% to 100%, cut tokens by more than half, and dropped task latency from 724 seconds to 141 seconds. A 2026 survey of harness engineering practices found that the same model running through a better-designed harness improved task success rates from 68% to 88%. The model did not change — the scaffold around it did.

The cost math is worth spelling out. Research on agent architectures puts baseline overhead at 311 tokens per tool call in an agentic loop. A five-step loop with two tool calls per step burns over 3,000 tokens before the model writes a single line of code. D-Engine skips that overhead entirely.

Why SEARCH/REPLACE Works

Forcing the LLM to output SEARCH/REPLACE blocks rather than arbitrary code introduces a grounding mechanism. Before the model can change anything, it must correctly identify what it is changing — it has to reproduce the original snippet. That reproduction step reduces hallucinations because the model is verifying its own understanding of the codebase before proposing a transformation.

Research from ACL 2026 on Search-and-Replace Infilling found this approach “harmonizes completion tasks with instruction-following priors of Chat LLMs” through structural grounding. The model plays to its strengths — pattern matching and targeted transformation — rather than holding an entire file in memory and regenerating it correctly. The SEARCH/REPLACE diff format also reduces output tokens significantly: for a 500-line file with a 10-line fix, the model returns the changed segments only, not the full file.

Getting Started

git clone https://github.com/corruchaga/D-Engine
cd D-Engine
npm install
cp .env.example .env
# Set LLM_BASE_URL, LLM_API_KEY, LLM_MODEL
npm run dev

D-Engine is LLM-agnostic. Point LLM_BASE_URL at any OpenAI-compatible endpoint — GPT-6.1 Sol, Gemini 4 Argon, Claude, or a local model via Ollama. The harness does not care which model generates the blocks; the TypeScript gate catches whatever the model gets wrong.

What It Cannot Do Yet

D-Engine cannot create new files — only edit existing ones (planned for a future release). The fuzzy matching tier is also the weakest point in the cascade; on heavily evolved codebases with inconsistent formatting, the 0.85 threshold can miss or misapply patches. The file selector introduces non-determinism at the input stage, which limits reproducibility on complex multi-file tasks. The benchmark covers 10 tasks — promising, but not statistically definitive.

These constraints are real. For greenfield work or tasks requiring new modules, D-Engine does not replace a full coding agent. It targets a specific, high-frequency workflow: making bounded, well-understood changes to existing TypeScript code.

The Bigger Picture

The AI coding tool market has spent two years competing on who can run the longest, most autonomous loops. D-Engine is a direct challenge to that direction. The assumption that more agency equals better results does not survive contact with the data. For a large class of everyday code edits — refactors, bug fixes, interface updates — a single-pass harness with a hard compilation gate is faster, cheaper, and safer than an open-ended loop. Developers burning API credits on multi-turn agents for routine changes should at least run the benchmark themselves.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *