-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathbenchmarks.html
More file actions
291 lines (266 loc) · 18.3 KB
/
Copy pathbenchmarks.html
File metadata and controls
291 lines (266 loc) · 18.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
<!DOCTYPE html>
<html lang="en" data-theme="dark">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>LoopGain Benchmark — 2,000 paired real-API trials, methodology pre-registered</title>
<meta name="description" content="LoopGain benchmark results: 92.8% cost reduction and ~15× wall-clock speedup vs max_iter=20 across 2,000 paired real-API trials. 10 cells, 6 framework adapters, 3 LLMs. Methodology pre-registered and locked 2026-05-21 before any confirmatory data was collected." />
<meta name="theme-color" content="#0A0B0D" />
<meta name="robots" content="index,follow" />
<link rel="icon" href="/favicon.ico" sizes="32x32" />
<link rel="icon" type="image/svg+xml" href="/favicon.svg" />
<link rel="apple-touch-icon" href="/apple-touch-icon.png" />
<link rel="canonical" href="https://loopgain.ai/benchmarks" />
<meta property="og:title" content="LoopGain Benchmark — 2,000 paired real-API trials" />
<meta property="og:description" content="92.8% cost reduction vs max_iter=20. ~15× wall-clock speedup. Output quality preserved on natural-distribution workloads and improved on engineered-failure ones. Pre-registered methodology; six analysis charts and raw data public." />
<meta property="og:url" content="https://loopgain.ai/benchmarks" />
<meta property="og:type" content="article" />
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link rel="preload" as="style" href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700&family=JetBrains+Mono:wght@400;500;600;700&display=swap" data-font-css />
<noscript><link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700&family=JetBrains+Mono:wght@400;500;600;700&display=swap" rel="stylesheet" /></noscript>
<script src="/fonts.js" defer></script>
<link rel="stylesheet" href="landing.css" />
<!-- Structured data — TechArticle (Schema.org) -->
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "TechArticle",
"headline": "LoopGain Benchmark — 2,000 paired real-API trials",
"description": "92.8% cost reduction vs max_iter=20 and ~15× wall-clock speedup across 2,000 paired real-API trials. 10 cells, 6 framework adapters, 3 LLMs. Methodology pre-registered 2026-05-21.",
"datePublished": "2026-05-25",
"url": "https://loopgain.ai/benchmarks",
"inLanguage": "en",
"author": { "@type": "Person", "name": "David Fitzsimmons" },
"publisher": {
"@type": "Organization",
"name": "LoopGain",
"url": "https://loopgain.ai",
"logo": "https://loopgain.ai/loopgain-lockup-dark.png"
},
"mainEntityOfPage": { "@type": "WebPage", "@id": "https://loopgain.ai/benchmarks" },
"isBasedOn": "https://github.com/loopgain-ai/loopgain-bench"
}
</script>
</head>
<body>
<!-- ─── nav ──────────────────────────────────────────────────────────── -->
<header class="nav" data-screen-label="nav">
<a class="nav-brand" href="/" aria-label="LoopGain home">
<img src="/loopgain-lockup-dark-sm.png" alt="LoopGain" class="brand-dark" style="height:24px; width:auto;" />
<img src="/loopgain-lockup-light-sm.png" alt="LoopGain" class="brand-light" style="height:24px; width:auto;" />
<span class="nav-version mono" data-lg-version>v0.5.2</span>
</a>
<nav class="nav-links" aria-label="primary">
<a href="/#bands">bands</a>
<a href="/#install">install</a>
<a href="/benchmarks">benchmark</a>
<a href="/#dashboard">dashboard</a>
<a href="/#pricing">pricing</a>
<a href="/blog/">blog</a>
<a href="https://github.com/loopgain-ai" rel="noopener">github ↗</a>
<button class="nav-theme" id="themeToggle" aria-label="Toggle theme" title="Toggle theme">
<svg width="16" height="16" viewBox="0 0 16 16" aria-hidden="true"><path d="M8 1.5v2M8 12.5v2M3.7 3.7l1.4 1.4M10.9 10.9l1.4 1.4M1.5 8h2M12.5 8h2M3.7 12.3l1.4-1.4M10.9 5.1l1.4-1.4" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" fill="none"/><circle cx="8" cy="8" r="3" stroke="currentColor" stroke-width="1.4" fill="none"/></svg>
</button>
<a class="nav-cta" href="https://dashboard.loopgain.ai" rel="noopener">open dashboard →</a>
</nav>
</header>
<main>
<article class="bench-page">
<span class="eyebrow mono"><span class="status-dot"></span> bench · locked 2026-05-25</span>
<h1>LoopGain Benchmark</h1>
<p class="bench-lede">
<strong>2,000 paired real-API trials. 10 cells across 6 framework adapters. Methodology pre-registered and locked 2026-05-21</strong> — before any of the confirmatory data was collected.
</p>
<p class="bench-note mono">Re-validated on loopgain v0.4.0 (carried unchanged into the current release); results landed 2026-06-03. The v0.4.0 classifier corrects a trajectory-label bug — a stuck loop now reads as STALLING, not OSCILLATING — and the corrected verdict terminates one iteration later, so the cost headline moved 93.5%→92.8%. Full re-validation note in the bench RESULTS.md.</p>
<!-- Aggregate hero (provenance: aggregate · all 2,000 trials). Replaced the
single-trial seed-34 two-run figure 2026-06-05: a single trial shown as two
independent runs can't isolate the stop rule from rollout variance (see
daves-kb/claim-provenance-public-vs-benchmark). The honest benchmark lead is
the aggregate distribution. Numbers pinned by loopgain-verify pub.bestdist_aggregate. -->
<figure class="bench-fig bench-bestdist"
aria-label="Across all 2,000 fixed-cap (max_iter=20) trials, 91.2% of loops reach their best output on the first iteration, but 8.8% land later — spread all the way out to iteration 20. No single static cap is both cheap and safe. LoopGain stops each loop at its own best: it keeps the best the loop reached on 97.2% of trajectories (2.8% false-stop) at a mean of 1.6 iterations, eliminating 97.2% of the wasted iterations at 92.8% lower cost.">
<div style="border:1px solid var(--border); border-radius:12px; padding:22px; background:var(--surf-1);">
<div class="mono" style="font-size:11px; color:var(--text-3); letter-spacing:0.04em; margin-bottom:18px;">
WHERE THE BEST OUTPUT LANDS · max_iter=20 · all 2,000 trials
</div>
<div style="display:grid; grid-template-columns:minmax(180px,0.7fr) 1.3fr; gap:32px; align-items:center;">
<div>
<div class="mono" style="font-size:52px; font-weight:600; line-height:1; letter-spacing:-0.03em; color:var(--band-conv);">91.2%</div>
<div style="font-size:13px; color:var(--text-2); margin-top:10px; line-height:1.45;">of loops hit their best output on the <strong>first</strong> iteration</div>
</div>
<div>
<div class="mono" style="font-size:11px; color:var(--text-3); margin-bottom:12px;">the other <span style="color:var(--text-1)">8.8% (176 loops)</span> land later — out to iteration 20:</div>
<div style="display:flex; flex-direction:column; gap:9px;">
<!-- 4 tail buckets, bar widths normalised among the tail (max=121) so they're visible -->
<div style="display:grid; grid-template-columns:70px 1fr 64px; gap:10px; align-items:center; font-size:11.5px;">
<span class="mono" style="color:var(--text-2)">iter 2–3</span>
<span style="height:11px; background:var(--surf-3); border-radius:2px; position:relative;"><span style="position:absolute; inset:0; width:100%; background:var(--band-stall); opacity:0.85; border-radius:2px;"></span></span>
<span class="mono" style="color:var(--text-1); text-align:right;">121</span>
</div>
<div style="display:grid; grid-template-columns:70px 1fr 64px; gap:10px; align-items:center; font-size:11.5px;">
<span class="mono" style="color:var(--text-2)">iter 4–5</span>
<span style="height:11px; background:var(--surf-3); border-radius:2px; position:relative;"><span style="position:absolute; inset:0; width:18.2%; background:var(--band-stall); opacity:0.85; border-radius:2px;"></span></span>
<span class="mono" style="color:var(--text-1); text-align:right;">22</span>
</div>
<div style="display:grid; grid-template-columns:70px 1fr 64px; gap:10px; align-items:center; font-size:11.5px;">
<span class="mono" style="color:var(--text-2)">iter 6–10</span>
<span style="height:11px; background:var(--surf-3); border-radius:2px; position:relative;"><span style="position:absolute; inset:0; width:16.5%; background:var(--band-stall); opacity:0.85; border-radius:2px;"></span></span>
<span class="mono" style="color:var(--text-1); text-align:right;">20</span>
</div>
<div style="display:grid; grid-template-columns:70px 1fr 64px; gap:10px; align-items:center; font-size:11.5px;">
<span class="mono" style="color:var(--text-2)">iter 11–20</span>
<span style="height:11px; background:var(--surf-3); border-radius:2px; position:relative;"><span style="position:absolute; inset:0; width:10.7%; background:var(--band-div); opacity:0.85; border-radius:2px;"></span></span>
<span class="mono" style="color:var(--text-1); text-align:right;">13</span>
</div>
</div>
</div>
</div>
<div style="margin-top:20px; padding-top:18px; border-top:1px solid var(--border); font-size:12.5px; color:var(--text-2); line-height:1.55;">
<strong>No static cap is both cheap and safe.</strong> A low cap clips the late-converging tail; a high cap grinds past best — and ships a measurably worse answer than the best it already had <strong>35.3%</strong> of the time. LoopGain stops each loop at its own best: it <strong style="color:var(--band-conv)">keeps the best on 97.2% of trajectories</strong> (2.8% false-stop), mean <strong>1.6 iterations</strong>, at <strong>92.8% less spend</strong>.
</div>
</div>
<figcaption>aggregate · iteration at which the best output appears · max_iter=20 condition · all 2,000 paired trials</figcaption>
</figure>
<span id="cost" class="bench-anchor" aria-hidden="true"></span>
<span id="latency" class="bench-anchor" aria-hidden="true"></span>
<span id="quality" class="bench-anchor" aria-hidden="true"></span>
<h2>The numbers</h2>
<p>
Across the full registered run — 10 cells × n=200 paired trials = 8,000 loop runs + 1,800 pairwise judge comparisons:
</p>
<div class="bench-table-wrap">
<table class="bench-table">
<thead>
<tr>
<th> </th>
<th>max_iter=5</th>
<th>max_iter=10</th>
<th class="col-emph">max_iter=20</th>
<th class="col-lg">LoopGain</th>
</tr>
</thead>
<tbody>
<tr>
<td>Total API spend</td>
<td>$6.83</td>
<td>$13.65</td>
<td class="col-emph">$27.05</td>
<td class="col-lg">$1.94</td>
</tr>
<tr>
<td>Median wall-clock per trial</td>
<td>7.2s</td>
<td>14.8s</td>
<td class="col-emph">30.9s</td>
<td class="col-lg">2.1s</td>
</tr>
<tr class="savings-row">
<td>Implied savings vs <code class="mono code-inline">max_iter=20</code></td>
<td>—</td>
<td>—</td>
<td>—</td>
<td class="savings-val">92.8% cost / 93.3% time</td>
</tr>
</tbody>
</table>
</div>
<p class="bench-note mono">Absolute wall-clock is environment- and concurrency-dependent and isn't a headline metric — this re-run ran on a less-contended machine than the 0.2.0 run, so the seconds are lower. The ~15× LG-vs-<code class="mono code-inline">max_iter=20</code> <em>ratio</em> is the stable claim; cost ratios are stable to ~1 pp run-to-run.</p>
<figure class="bench-fig">
<img
src="https://raw.githubusercontent.com/loopgain-ai/loopgain-bench/main/data/results/charts/cost_by_condition.png"
alt="Bar chart: total API spend by condition. max_iter=5 = $6.83, max_iter=10 = $13.65, max_iter=20 = $27.05, LoopGain = $1.94."
loading="lazy" />
<figcaption>cost_by_condition.png · total API spend across the bench, by condition</figcaption>
</figure>
<ul>
<li><strong>92.8% cost reduction</strong> vs <code class="mono code-inline">max_iter=20</code></li>
<li><strong>~15× wall-clock speedup</strong> (30.9s → 2.1s median per trial)</li>
<li><strong>Quality preserved</strong> on natural-distribution workloads (W1–W2 judge winrate 0.55–0.63 with CI excluding null on most cells; W3 cells tie-dominated as preservation-by-construction since both LoopGain and <code class="mono code-inline">max_iter=20</code> produce the same correct tool call ~90% of the time)</li>
<li><strong>Quality improved</strong> on engineered-failure workloads (W5 winrate 0.92–0.95 across three adapters via best-so-far rollback)</li>
<li><strong>Aggregate</strong> weighted-average pairwise preference for LoopGain across 1,800 judge comparisons: <strong>0.678</strong> — over two-thirds</li>
<li><strong>Zero of six kill criteria fired</strong></li>
</ul>
<h2>See it live</h2>
<div class="bench-cta">
<a class="bc-link" href="https://dashboard.loopgain.ai/benchmark" rel="noopener">Open the bench data in the LoopGain dashboard →</a>
<p>
The bench tenant's 2,000 trials are visible in the actual product dashboard — the same UI a customer would see, populated with the canonical benchmark data. Read-only public view; sign up free to instrument your own loops.
</p>
</div>
<h2>Honest disclosures</h2>
<p>One pre-registered floor was missed without firing a kill criterion. It's surfaced in the writeup:</p>
<ul>
<li><strong>H-FRAMEWORK-PARITY</strong> W2 spread: 5.5 pp observed vs ≤5 pp predicted (kill threshold: >15 pp)</li>
</ul>
<p class="bench-note mono">Under the original 0.2.0 run, H-EARLYWARN missed at 2 iterations and the widest parity spread was W1 at 5.8 pp; the corrected 0.4.0 classifier flags STALLING earlier, so median lead time now meets the ≥3-iteration floor and the widest spread moved to W2 at 5.5 pp.</p>
<p>
Seven pre-data amendments to the methodology are preserved in <a href="https://github.com/loopgain-ai/loopgain-bench/blob/main/BENCH_PROTOCOL.md" rel="noopener">BENCH_PROTOCOL.md</a> with their full rationale. Predicted floors and kill criteria were never changed once data started landing.
</p>
<p>
The bench harness itself had two non-trivial bugs caught and fixed during the run (signal-handling under concurrency; thread-pool shutdown semantics) — both forensics documented honestly in <a href="https://github.com/loopgain-ai/loopgain-bench/blob/main/LESSONS.md" rel="noopener">LESSONS.md</a>. The data here is from the post-fix, n=200, tripwire-clean run.
</p>
<h2>Reproduce it yourself</h2>
<p>
The bench, its methodology, the raw data (9.4 MB of JSONL across 10 cells + 9 judge runs), all six analysis charts, the engineering forensics, and all seven pre-data amendments are public:
</p>
<p>
<strong><a href="https://github.com/loopgain-ai/loopgain-bench" rel="noopener">github.com/loopgain-ai/loopgain-bench</a></strong>
</p>
<pre class="bench-code"><span class="tk-com">$</span> git clone https://github.com/loopgain-ai/loopgain-bench
<span class="tk-com">$</span> cd loopgain-bench
<span class="tk-com">$</span> make install-dev
<span class="tk-com">$</span> make bench <span class="tk-com"># ~$50, ~4-8h on a single Mac</span>
<span class="tk-com">$</span> make judge <span class="tk-com"># ~$1-2</span>
<span class="tk-com">$</span> make analyze <span class="tk-com"># six tables + six charts</span></pre>
<section class="bench-try">
<h2>Try LoopGain</h2>
<p>
<code class="mono code-inline">pip install loopgain</code> — or <code class="mono code-inline">pip install 'loopgain[<your framework>]'</code> for adapter extras (LangGraph, CrewAI, AutoGen, LangChain, OpenAI Agents SDK, Claude Agent SDK).
</p>
<div class="bench-try-actions">
<div class="copy-box" data-copy="pip install loopgain" role="button" tabindex="0">
<span class="copy-prompt mono">$</span>
<code class="mono">pip install loopgain</code>
<span class="copy-hint mono">copy</span>
</div>
<a class="btn btn-amber" href="https://dashboard.loopgain.ai" rel="noopener">
open dashboard
<svg width="12" height="12" viewBox="0 0 12 12" aria-hidden="true"><path d="M3 9 L9 3 M5 3 H9 V7" stroke="currentColor" stroke-width="1.4" fill="none" stroke-linecap="round"/></svg>
</a>
</div>
<p class="bench-try-pilot">
<strong>3 design-partner slots open for the next 30 days.</strong> Free 30-day pilot, direct founder support. Email <a href="mailto:hello@loopgain.ai?subject=Design-partner%20pilot">hello@loopgain.ai</a>.
</p>
</section>
<p class="bench-back">← <a href="/">back to loopgain.ai</a></p>
</article>
</main>
<!-- ─── footer ───────────────────────────────────────────────────────── -->
<footer class="foot" data-screen-label="footer">
<div class="foot-row">
<a class="foot-brand" href="/" aria-label="LoopGain home">
<img src="/loopgain-lockup-dark-sm.png" alt="LoopGain" class="brand-dark" style="height:20px; width:auto;" />
<img src="/loopgain-lockup-light-sm.png" alt="LoopGain" class="brand-light" style="height:20px; width:auto;" />
</a>
<nav class="foot-links mono" aria-label="footer">
<a href="https://github.com/loopgain-ai" rel="noopener">github</a>
<a href="https://pypi.org/project/loopgain" rel="noopener">pypi</a>
<a href="https://dashboard.loopgain.ai" rel="noopener">dashboard</a>
<a href="/blog/">blog</a>
<a href="/benchmarks">benchmark</a>
<a href="/privacy">privacy</a>
<a href="/terms">terms</a>
<a href="mailto:hello@loopgain.ai">hello@loopgain.ai</a>
</nav>
</div>
<div class="foot-row foot-row-fine mono">
<span>© 2026 LoopGain Technologies Inc. · Apache-2.0</span>
<span class="dot-sep"></span>
<span>Cost control for AI agent loops</span>
</div>
</footer>
<script src="landing.js"></script>
<script src="version-sync.js"></script>
</body>
</html>