Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Evaluation protocol

Status: draft. Not yet pre-registered. The decision rule in §6 becomes binding when it is committed and its commit hash is recorded in docs/research/evaluation-preregistration.md. That file does not exist; until it does, the rule is a proposal and may change freely. After that, it may not.

Scope changed 2026-09-20, contents unchanged. This protocol no longer decides whether the language project exists. It is now evidence for specific usability and performance claims, and it is reproduced here exactly as registered-in-draft: the arms, the metrics, the thresholds in §6 (point R_C ≤ 0.80 and upper CB < 1.00), the cost policy, and the void-rate rule are untouched. Nothing has been relaxed to fit the new direction, and no earlier result has been reinterpreted.

The decision record that changed this scope (direction.md) was retired on 2026-09-21 once each of its standing decisions had an owner; Git retains it. The decision itself, restated here because this is the document it governs:

The language is the primary track. This protocol is evidence for claims, not a stop condition for the project. Arm A2 keeps its role — if improved TypeScript tooling captures most of a measured benefit, that is a finding about that claim and it stays in the record. The claims it is now evidence for: that agents spend less effort per accepted change in Nazm than in TypeScript with comparable tooling; that structured diagnostics and a canonical formatter reduce rework; and eventually the runtime claims in ../performance.md, which are a separate experiment and must never be collapsed into one score with this one.

A2 keeps its role. If improved tooling on an existing language captures most of a measured benefit, that is a finding about the claim under test and it stays in the record. It is no longer a stop condition.


1. The question

Does a new language reduce the cost of building and maintaining backend services with coding agents, over and above what better tooling on an existing language achieves?

That last clause is the whole experiment. Without it, any improvement gets attributed to the language by default, which is the mistake this protocol exists to prevent.

2. Arms

ArmContentsIsolates
A1TypeScript, baseline tooling (tsc, eslint, prettier, stock agent loop)starting point
A2TypeScript + improved agent tooling: structured diagnostics with machine-applicable fixits, canonical formatter, repeat-edit detection, explicit completion rulesthe tooling contribution
A3Nazm + tooling of comparable capabilitythe additional language contribution

Go is run as a runtime performance and resource reference only. It is not an arm: runtime performance and agent development efficiency are different questions and are reported separately, never combined into one score.

The primary comparison is A3 against A2. A3-against-A1 is reported but decides nothing.

3. Tasks

Backend (primary domain)

TaskWhat it tests
Validated CRUD APIBasic language usability, JSON, diagnostics
Transactional inventory reservationCorrectness across concurrent requests
Tenant-isolated APIAuthorization and data-access rules
Idempotent webhook handlerRetries, duplicate delivery, persistence
Cancellable background jobResource cleanup and task lifecycle
Feature change in an existing serviceAgent comprehension and regression avoidance

Each is a family, instantiated into multiple concrete tasks. Target 30–50 total.

A smaller ETL and CLI sample runs alongside to test whether another domain shows a stronger advantage. Backend remains the working choice; it is revised if the evidence warrants, and that revision is reported.

The maintenance task matters most and is the one most benchmarks omit. Short greenfield algorithm problems are the easiest thing to measure and the least like the work.

Three sets, and a freeze

Hiding acceptance tests from the agent is necessary but not sufficient. If Nazm, its wrappers, and its diagnostics are tuned against the same tasks for a year, the final benchmark rewards specialisation to that set rather than language quality.

SetPermitted use
DevelopmentLanguage design, diagnostics, wrappers, debugging. Opened freely
CalibrationVariance estimation, power analysis, protocol rehearsal
HoldoutGate evaluation only, after implementation freeze. Never opened before

Holdout isolation is operational, not type-level. The harness refuses to hand back a holdout fixture without a Freeze, and a fabricated Freeze fails when verified against the repository — but that guards the harness API, not the filesystem. Anything with read access to the corpus directory can bypass it, and repository history, logs, and CI artifacts can all leak fixture content. For the gate run the holdout corpus must live outside the agent’s accessible environment entirely: a separate host or credential-gated store, fetched after the freeze and never mounted where the agent can reach it.

The freeze is a git tag. Everything under evaluation — compiler, tooling, prompts, wrappers — is fixed at that tag, and the holdout run uses only that tag.

4. Fairness conditions

Held equivalent: task requirements, evaluation rules, and the independent acceptance tests, which the agent never sees and which execute outside the agent’s writable environment.

Allowed to differ, and recorded: idiomatic implementation style, language-specific diagnostics. Forcing one language to imitate another’s idioms measures the wrong thing.

Comparable starting codebases. For maintenance tasks the TypeScript and Nazm services must be functionally equivalent and of comparable quality. A crufty TypeScript repo against a freshly simplified Nazm one measures codebase quality, not language.

Aligned semantic assistance. If Nazm ships reserveInventorySafely() while TypeScript must hand-roll the transaction, the experiment measures library abstraction. The helper surface available to each arm is inventoried, aligned, and published with the results.

“Same budget” is pinned down. Dollar, token, turn, and wall-clock caps are four different constraints. The pre-registration names which bind and fixes model versions, reasoning-effort settings, tool access, dependency availability, and context configuration.

Arms are interleaved within each trial block, so drifting service conditions do not land disproportionately on one arm.

5. Metrics

Primary:

cost per accepted solution = total cost of all attempts / number of accepted solutions

This metric is perverse alone and is never reported alone. An arm accepting 40/100 at $40 scores $1.00, beating an arm accepting 90/100 at $100 scoring $1.11 — while failing most of the work. Always reported with it:

  • acceptance rate (absolute, and delta against A2)
  • per-task-category breakdown
  • wall clock
  • regressions introduced (maintenance tasks)
  • human interventions required
  • unsuccessful repair cycles
  • void count and void rate, per arm — evaluator faults stay outside the acceptance-rate denominator under the registered policy, but that licence has a registered limit (max_void_rate). Past it the comparison is no longer between two arms, it is between one arm and a partly-broken measurement of another, and the run is repeated or invalidated rather than averaged

Cost is three separate lines, never silently summed into one: model charges · execution infrastructure · logged human intervention.

All attempts count, including failures. Post-success activity is categorised (cosmetic / revert / performance / security / edge case) to explain where cost went — the categories are analysis labels, never a licence to exclude activity from the denominator. Some reverts repair real regressions.

6. The decision rule

Not yet binding. See the status note at the top.

Let p be acceptance rate and C cost per accepted solution:

Δp  = p₃ − p₂
R_C = C₃ / C₂

Uncertainty is estimated on the difference and the ratio themselves, not by checking whether two separate confidence intervals overlap — those are different questions and the second is not a test.

Repeated runs of the same task are not independent samples. The Tokenmaxxing study found problem identity explains 73–97% of cost variance. Task enters the model as a random effect, with model (the LLM) as a second grouping factor. Aggregation across models and categories is fixed in the pre-registration, not chosen after seeing results.

Cost policy

Two defensible standards exist. The second is chosen:

PolicyRequirementStatus
Evidence of ≥20% savingsupper confidence bound of R_C ≤ 0.80rejected — too strong for a 30–50 task benchmark with agent-level noise; would reject a real effect
Estimated 20% savings, with evidence of some savingspoint R_C ≤ 0.80 and upper CB of R_C < 1.00chosen — appropriate for a research gate; trial count is powered for it

Continue — requires all

ConditionThreshold
Costpoint R_C ≤ 0.80 and upper CB of R_C < 1.00
Non-inferioritylower CB of Δp ≥ −5 percentage points
Absolute usability floorA3 acceptance ≥ 60% absolute — relative non-inferiority alone permits both arms to be unusable
No category collapseno task category where A3 acceptance < 70% of A2’s
Maintenance holdson feature change in an existing service: acceptance not lower and regression count not higher than A2. Cost reported but not gating — the maintenance sample is small

Redesign

Cost fails but every acceptance condition holds, or the gain sits in a single category. Identify which mechanism underperformed, take one targeted iteration, re-gate under a fresh pre-registration.

Stop

Any acceptance condition fails and cannot be attributed to a fixable defect in the harness, or the result is inconclusive.

Inconclusive means stop. The trial count is fixed in advance by power analysis at the end of Phase 3. There is no “run more trials if it is trending our way” clause: that requires a prespecified sequential-analysis boundary, and without one it is optional stopping. Any further experiment is a separate, separately pre-registered evaluation.

Tooling capture — reported, never a stop rule

Capture = (C₁ − C₂) / (C₁ − C₃), the fraction of the total improvement attributable to tooling alone.

An earlier draft made “capture ≥ 80% ⇒ stop” an automatic rule. It contradicts the continue rule. At C₁=$100, C₂=$20, C₃=$15: Nazm is 25% cheaper than A2, so continue fires; capture is 80/85 ≈ 94%, so stop also fires. Both cannot bind.

A large tooling win does not make an additional language win worthless. Capture is published as a descriptive figure. The investment decision rests on A3’s incremental benefit over A2, its reliability, and the cost of building and maintaining the language — the last being a judgement, made explicitly and recorded, not a threshold.

Chronology, not just consistency

A registration that merely agrees with the documentation proves nothing: both files can be written in one commit after the results have been examined. So the registration is bound two ways, and neither alone is sufficient.

Git ancestry — every attempt record names the registration commit, which must be a strict ancestor of the run commit. This establishes ordering within repository history. It does not establish that the registration preceded the experiment: history is authored, and someone can run first, then create both commits in the required order.

A preflight record — before any agent is launched, the runner verifies the registration and freeze, captures the run commit, task-bundle digest, configuration, and dependency digest, refuses to proceed from a dirty tree (or captures the diff and marks the run unpublishable), and writes a start record. This ties the registration to a specific execution rather than to a commit graph.

Neither is an independent timestamp, and this protocol does not claim one. If publication requires chronology verifiable by a third party, the registration must also be recorded externally — a push whose time the forge records, or a timestamping service — and that is an additional step, not something the repository provides.

A registration binds the protocol and analysis rules (its own content), the evaluated implementation (the freeze tag), the task sets (a corpus digest), and its version. A later protocol change creates a new registration; earlier runs keep pointing at the one they ran under and stay interpretable under the rules that applied to them.

7. What the gate establishes

It is the primary comparative evaluation of Nazm’s agent-development benefit. It is not proof of the project. It establishes whether Nazm adds value for these tasks, these models, and these configurations. Nothing more.

It does not establish that Nazm is a good language, that the design is sound, that the results generalise to other domains, or that the advantage persists as either arm matures.

8. Harness validation

Validated by construction, not by agreement with published results.

Known-answer tests must confirm correct handling of:

  • each outcome class — passing, wrong answer, compile failure, crash, timeout
  • an agent that modifies the tests, or invokes the wrong executable
  • missing token-usage records, retries, interrupted runs
  • state leaking between attempts
  • a solution that passes the visible tests but fails the independent acceptance tests

Missing cost data invalidates or explicitly qualifies a result. It must never silently become zero cost — that single failure mode can manufacture an arbitrary winner.

Then report what the runs show, including results that contradict the published direction. A different outcome can come from different models, tasks, or conditions; it is not evidence the harness is broken. An earlier draft of the plan had this backwards and said a non-reproducing result meant the harness was wrong. That assumes the conclusion.

9. Prior results this protocol is calibrated against

Cited as context and as a sanity check on the instrument — not as thresholds, and not as findings this protocol expects to reproduce.

Figures from that paper are cell-specific and are not general properties of coding agents: the “92% of failures are compile errors” result is the Gemma–OCaml cell, and the “3.6× post-success tokens” result is Gemma–Rust medians.