|
| 1 | +--- |
| 2 | +title: Legacy Cross-model Results |
| 3 | +description: Earlier four-model NetOpsBench comparison retained for within-snapshot analysis. |
| 4 | +--- |
| 5 | + |
| 6 | +This snapshot predates the NetOpsBench v0.2.0 release rerun. Its four models were evaluated under the same earlier contract, so it remains useful for within-snapshot comparison. Do not combine these values with the current [v0.2.0 Results](/docs/run-benchmarks/results) as if both snapshots used the same runtime and detector contract. |
| 7 | + |
| 8 | +## Experiment scope |
| 9 | + |
| 10 | +| Dimension | Values | |
| 11 | +|---|---| |
| 12 | +| Models | Kimi K2.6, DeepSeek V4 Pro, OpenAI GPT-5.5, MiniMax M3 | |
| 13 | +| Topology scales | XS, Small, Medium, Large | |
| 14 | +| Fault types | 12 canonical types across link, routing, impairment, system, and ACL categories | |
| 15 | +| Quality metrics | Verdict F1-score, device localization, interface localization, composite score | |
| 16 | +| Efficiency metrics | Diagnosis time, tool calls, input/output tokens | |
| 17 | + |
| 18 | +## Cross-scale snapshot |
| 19 | + |
| 20 | +Each cell is `XS → Large`. |
| 21 | + |
| 22 | +| Model | Verdict F1 (%) | Device Loc (%) | Interface Loc (%) | Composite (%) | Avg Time (s) | Avg Tools | Input Tokens (K) | |
| 23 | +|---|---:|---:|---:|---:|---:|---:|---:| |
| 24 | +| Kimi K2.6 | 76.2 → 84.7 | 66.7 → 75.0 | 71.4 → 71.4 | 64.3 → 74.0 | 272.0 → 399.5 | 36.5 → 38.8 | 367.6 → 741.1 | |
| 25 | +| DeepSeek V4 Pro | 100.0 → 97.9 | 83.3 → 91.7 | 57.1 → 57.1 | 78.6 → 83.7 | 83.1 → 83.1 | 24.9 → 20.7 | 247.6 → 477.3 | |
| 26 | +| OpenAI GPT-5.5 | 95.7 → 95.8 | 83.3 → 79.2 | 57.1 → 21.4 | 75.0 → 60.6 | 71.9 → 85.0 | 17.9 → 15.9 | 85.4 → 153.7 | |
| 27 | +| MiniMax M3 | 80.0 → 82.9 | 50.0 → 62.5 | 71.4 → 42.9 | 60.7 → 58.7 | 214.3 → 230.3 | 26.4 → 24.0 | 133.3 → 316.1 | |
| 28 | + |
| 29 | +Verdict classification was generally easier than precise localization. Interface localization was the weakest quality metric on larger topologies. |
| 30 | + |
| 31 | +## Quality metrics |
| 32 | + |
| 33 | + |
| 34 | + |
| 35 | + |
| 36 | + |
| 37 | + |
| 38 | + |
| 39 | + |
| 40 | + |
| 41 | +## Runtime cost |
| 42 | + |
| 43 | + |
| 44 | + |
| 45 | + |
| 46 | + |
| 47 | + |
| 48 | + |
| 49 | + |
| 50 | + |
| 51 | +Compare cost with localization quality, not verdict quality alone. Higher tool or token use does not automatically improve device or interface precision. |
0 commit comments