<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>AHD Eval Runs</title>
  <subtitle>Every AHD raw-versus-compiled eval run, dated, versioned, with manifests and per-cell counts.</subtitle>
  <link rel="self" type="application/atom+xml" href="https://ahd.adastra.computer/evals/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals"/>
  <id>https://ahd.adastra.computer/evals/feed.xml</id>
  <updated>2026-08-10T00:00:00Z</updated>
  <rights>FSL-1.1-Apache-2.0</rights>
  <generator uri="https://ahd.adastra.computer">AHD site build</generator>
  <entry>
    <id>https://ahd.adastra.computer/evals/weekly-swiss-n30</id>
    <title>Eight runs · Same brief</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/weekly-swiss-n30"/>
    <updated>2026-08-10T00:00:00Z</updated>
    <published>2026-08-10T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">The weekly programme read as a series rather than eight bulletins, covering 9 June to 10 August 2026. gpt-oss-120b and mistral-small-3.1 reduce every run and by a consistent amount. gemma-4 reduces every run by an amount that swings twenty points. llama-4-scout sits against zero. qwen3-30b changes sign six times in eight runs, so no claim about its direction survives the series. New runs update this page rather than adding another.</summary>
    <category term="Weekly series · CF OSS n=30"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-06-22-swiss-n30</id>
    <title>Three weeks · One split</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-06-22-swiss-n30"/>
    <updated>2026-06-22T00:00:00Z</updated>
    <published>2026-06-22T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">Third consecutive weekly run. Gemma 57.2%, mistral 64.6% and gpt-oss 72.7% reduce under the compiled prompt; llama-4-scout stays flat at 1.6%. qwen3 swings back to -7.0% after +7.1% the week before, straddling zero across three runs (-3.4, +7.1, -7.0). The reducing band and the llama trade hold; qwen is the lone unstable cell.</summary>
    <category term="Weekly · CF OSS n=30"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-06-15-swiss-n30</id>
    <title>The split holds · Week two</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-06-15-swiss-n30"/>
    <updated>2026-06-15T00:00:00Z</updated>
    <published>2026-06-15T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">Second weekly run. The 9 June split reproduces: gemma 53.3%, mistral 66.0% and gpt-oss 73.0% reduce under the compiled prompt; llama-4-scout flat at 0.0% and qwen3 at 7.1%. llama again trades named-grid and type-pairing for line-height and radius rather than reducing. Repeating two weeks running makes this a stable pattern, not a one-off.</summary>
    <category term="Weekly · CF OSS n=30"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-06-09-swiss-n30</id>
    <title>Five models · Three reduce · Two can&apos;t follow</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-06-09-swiss-n30"/>
    <updated>2026-06-09T00:00:00Z</updated>
    <published>2026-06-09T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">First run on the automated weekly cadence. Five Cloudflare Workers AI open-source models, n=30, source-linter only. Three reduced cleanly under the compiled prompt: gpt-oss-120b 72.6%, mistral-small-3.1 68.6% and gemma-4 53.8%. Two stayed flat: llama-4-scout at 0% and qwen3 at -3.4%. The flat cells reflect the models, not the tooling. llama-4-scout trades tells, dropping require-named-grid and require-type-pairing but introducing a single uniform line-height and radius that the token tells it to vary. The linter correctly separates a model that follows per-size guidance from one that cannot.</summary>
    <category term="Weekly · CF OSS n=30"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-04-24-post-digital-green-n30</id>
    <title>Eleven models · Same brief · Different token</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-04-24-post-digital-green-n30"/>
    <updated>2026-04-24T00:00:00Z</updated>
    <published>2026-04-24T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">The different-token-same-brief triangulation queued by the 22 April report. Eight of eleven cells regress under the compiled prompt. The compiler is not at fault: it transmits the token faithfully. The regressions come from lint rules that assumed the editorial defaults this token rejects, so they penalised output that followed the token. gpt-5.5 lands as the cleanest raw frontier baseline measured to date at 1.03 tells per page. Token-aware linting is the next engineering step.</summary>
    <category term="Cross-provider n=30 · post-digital-green"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-04-22-swiss-n30</id>
    <title>Ten models · One brief · Thirty samples each</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-04-22-swiss-n30"/>
    <updated>2026-04-22T00:00:00Z</updated>
    <published>2026-04-22T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">Ten models, n=30 per cell, 600 samples. Eight of ten cells showed positive reductions under the compiled prompt, one flat, one regression. Best: gpt-oss-120b at 78.1% fewer tells. Three frontier cells via subscription CLIs (Claude Code, Codex, Gemini CLI), seven OSS via Cloudflare Workers AI. Wilson interval tightens from roughly +/-35% at n=5 to roughly +/-18% at n=30.</summary>
    <category term="Cross-provider n=30"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-04-21-swiss-cross</id>
    <title>Seven models across four providers</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-04-21-swiss-cross"/>
    <updated>2026-04-21T00:00:00Z</updated>
    <published>2026-04-21T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">Four positive reductions, one inconclusive, two regressions. Llama 3.3&apos;s regression reproduces across Cloudflare and Hugging Face, turning a single-cell finding into a cross-provider result.</summary>
    <category term="Cross-provider n=5"/>
  </entry>
  <entry>
    <id>https://ahd.adastra.computer/evals/2026-04-21-swiss</id>
    <title>Five models, n=5, zero errors</title>
    <link rel="alternate" type="text/html" href="https://ahd.adastra.computer/evals/2026-04-21-swiss"/>
    <updated>2026-04-21T00:00:00Z</updated>
    <published>2026-04-21T00:00:00Z</published>
    <author>
      <name>Ad Astra Computing Inc</name>
      <uri>https://adastracomputing.com</uri>
    </author>
    <summary type="text">Claude Opus (Anthropic API) plus four OSS models on Cloudflare Workers AI. Claude dropped to zero tells compiled; Llama 3.3 70B regressed. Full per-model and per-tell breakdown with every attempted-vs-scored count published.</summary>
    <category term="Five-model n=5"/>
  </entry>
</feed>
