|
2 | 2 |
|
3 | 3 | > Auto-generated by `benchmarks/`. Run `uv run python -m benchmarks.bench_X` to refresh. |
4 | 4 |
|
5 | | -Current status: 🚧 awaiting TraCSS data extraction (20.73 GB still downloading). |
| 5 | +Last updated: 2026-05-21 (first real-data calibration run). |
| 6 | + |
| 7 | +## Status |
| 8 | + |
| 9 | +**Dataset extracted, pipeline calibrated on real Aerospace IVV OCMs.** |
| 10 | + |
| 11 | +| What | Status | |
| 12 | +|---|---| |
| 13 | +| TraCSS 20.73 GB dataset | ✅ Downloaded + extracted to `data/tracss/` | |
| 14 | +| Answer keys (913,330 spherical + 283,595 SFSH conjunctions) | ✅ Decompressed and loaded | |
| 15 | +| OCM parser | ✅ Calibrated on real Aerospace IVV files (CCSDS 502.0-B-3 KVN) | |
| 16 | +| End-to-end pipeline | ✅ Runs without errors on real OCMs | |
| 17 | +| First benchmark vs answer key | ✅ See B4 below | |
| 18 | + |
| 19 | +--- |
6 | 20 |
|
7 | 21 | ## B1 — Propagation throughput |
8 | 22 |
|
9 | 23 | | Metric | Target | Current | |
10 | 24 | |---|---|---| |
11 | 25 | | Full Starlink (9341 sats) × 1000 steps | <10 ms on A100 | Pending GPU access | |
12 | | -| Local CPU smoke (1000 × 100) | — | Run `uv run python -m benchmarks.bench_propagate --smoke` | |
| 26 | +| Local CPU smoke (100 × 10) | — | ✅ 12.1 ms (M-series Mac, JIT warm-cache) | |
13 | 27 |
|
14 | 28 | ## B2 — Screening throughput |
15 | 29 |
|
16 | 30 | | Metric | Target | Current | |
17 | 31 | |---|---|---| |
18 | | -| Octree vs naive O(N²) at N=10K | >1000× speedup | Pending full run | |
| 32 | +| Octree vs naive O(N²) at N=10K | >1000× speedup | Pending full GPU run | |
| 33 | +| Smoke (100 objects, 10 km radius) | — | ✅ Octree 2.3 ms vs Naive 2.9 ms | |
19 | 34 |
|
20 | 35 | ## B3 — Pc method correctness |
21 | 36 |
|
22 | 37 | | Method | Target (vs Monte Carlo) | Current | |
23 | 38 | |---|---|---| |
24 | | -| Alfano 2004 | median rel err < 0.05 | Pending | |
25 | | -| Chan | median rel err < 0.05 | Pending | |
26 | | -| Foster | matches MC 89% (Auman 2025) | Pending | |
27 | | -| Patera | matches Alfano within 5% | Pending | |
| 39 | +| Alfano 2004 (primary) | median rel err < 0.05 | ✅ Method implemented, tests passing | |
| 40 | +| Chan | median rel err < 0.05 | ✅ Tests passing | |
| 41 | +| Foster | matches MC ~89% (Auman 2025) | ✅ Tests passing | |
| 42 | +| Patera | matches Alfano within 5% | ✅ Tests passing | |
28 | 43 |
|
29 | | -## B4 — TraCSS Spherical screening |
| 44 | +Full cross-validation against published NASA CARA fixtures pending fixture import. |
30 | 45 |
|
31 | | -| Metric | Target | Current | |
32 | | -|---|---|---| |
33 | | -| Recall vs answer key | ≥99% | Pending extraction | |
34 | | -| Precision vs answer key | ≥99% | Pending extraction | |
35 | | -| Median TCA diff | <5 sec | Pending | |
| 46 | +## B4 — TraCSS Spherical screening (first real-data run) |
| 47 | + |
| 48 | +**Test subset: 79 OCMs from the Aerospace IVV dataset.** Within the answer key |
| 49 | +restricted to those 79 sat IDs: **16 distinct conjunction pairs / 26 total conjunctions.** |
| 50 | + |
| 51 | +| Metric | Target | v1 (no fixes) | v2 (local minima) | **v3 (+ TCA refine)** | |
| 52 | +|---|---|---|---|---| |
| 53 | +| Distinct pairs found | 16 | 2 | 3 | **10** | |
| 54 | +| Pair-level recall | ≥99% | 12.5% | 18.7% | **62.5%** | |
| 55 | +| Pair-level precision | ≥99% | 100% | 100% | **100%** | |
| 56 | +| Conjunction-level recall | ≥99% | 7.7% | 15.4% | **42.3%** | |
| 57 | +| Wall clock (79 OCMs, M5 Pro CPU) | — | 196s | 325s | 487s | |
| 58 | + |
| 59 | +**Interpretation:** |
| 60 | +- ✅ Zero false positives across all runs — when we flag a conjunction, it's real. |
| 61 | +- ✅ TCA refinement (golden-section search ± 2 min) lifted recall from 12.5% to 62.5%. |
| 62 | +- ⚠️ Six pairs still missed — likely fast-flyby pairs that need an even wider coarse screening radius. Iteration planned. |
| 63 | + |
| 64 | +Missed pairs from this subset (all flagged in answer key but not by our pipeline): |
| 65 | +- 15755-51103, 53072-95034, 53700-95034, 58899-95343, 95236-95237, 99000-99002 |
| 66 | + |
| 67 | +Path to ≥99%: (a) widen coarse screening radius to 500+ km, (b) use 60s sampling, (c) implement velocity-aware "swept volume" pre-check between samples. |
36 | 68 |
|
37 | 69 | ## B5 — TraCSS SFSH screening |
38 | 70 |
|
39 | | -Same metrics as B4 with per-object rectangular volumes. Pending extraction. |
| 71 | +Pending — same harness, just plug in SFSH per-object volumes. Will run after B4 recall hits ≥95%. |
40 | 72 |
|
41 | 73 | ## B6 — End-to-end wall clock |
42 | 74 |
|
43 | 75 | | Metric | Target | Current | |
44 | 76 | |---|---|---| |
45 | | -| 30K objects, screen + Pc + emit CDMs | <30 sec on A100 | Pending GPU run | |
46 | | -| 100 objects local | — | Run `uv run python -m benchmarks.bench_end2end --smoke` | |
| 77 | +| 30K objects, screen + Pc + emit CDMs | <30 sec on A100 | 487s (79 OCMs, M5 Pro CPU) | |
| 78 | +| Per-object cost extrapolation | — | ~6 sec/OCM on CPU (single-threaded) | |
| 79 | + |
| 80 | +The CPU run is single-threaded Python loops. Path to <30s at 30K: (a) `jax.vmap` the per-time-step interp, (b) GPU acceleration via Modal, (c) parallel pair-trace minimization. |
47 | 81 |
|
48 | 82 | ## B7 — Maneuver opt quality |
49 | 83 |
|
50 | 84 | | Metric | Target | Current | |
51 | 85 | |---|---|---| |
52 | | -| Δv vs greedy heuristic | 2× lower | Pending | |
| 86 | +| Δv vs greedy heuristic | 2× lower | ✅ Optimizer ships; benchmark against greedy pending | |
53 | 87 |
|
54 | 88 | ## B8 — Agent latency |
55 | 89 |
|
56 | 90 | | Metric | Target | Current | |
57 | 91 | |---|---|---| |
58 | | -| p50 | <5 sec | Pending API key | |
59 | | -| p95 | <15 sec | Pending API key | |
| 92 | +| p50 | <5 sec | Pending Anthropic API key | |
| 93 | +| p95 | <15 sec | Pending | |
| 94 | + |
| 95 | +Agent code is structurally complete; stub mode returns deterministic responses for testing without API. |
60 | 96 |
|
61 | 97 | ## B9 — Agent correctness |
62 | 98 |
|
63 | | -| Metric | Target | Current | |
64 | | -|---|---|---| |
65 | | -| Answers match hand-computed reference | ≥95% on 50 hand-crafted prompts | Pending | |
| 99 | +Pending — needs 50 hand-crafted reference queries to grade against. |
| 100 | + |
| 101 | +--- |
| 102 | + |
| 103 | +## How to reproduce |
| 104 | + |
| 105 | +```bash |
| 106 | +# 1. Get TraCSS access form approved (see README) |
| 107 | +# 2. Download and extract: |
| 108 | +cd data/tracss |
| 109 | +gunzip -k *.csv.gz |
| 110 | +tar -xzf AerospaceIVVDataset_20251009a.tar.gz |
| 111 | + |
| 112 | +# 3. Run subset benchmark: |
| 113 | +uv run python -c " |
| 114 | +from skyshield.eval.tracss_runner import run_tracss_screening, write_cdm_csv |
| 115 | +result = run_tracss_screening( |
| 116 | + 'data/tracss/AerospaceIVVDataset_20251009', |
| 117 | + pattern='*.ocm', |
| 118 | + mode='spherical', |
| 119 | + screening_radius_km=100.0, |
| 120 | + time_step_seconds=120.0, |
| 121 | +) |
| 122 | +filtered = [c for c in result.conjunctions if c.min_range <= 10.0] |
| 123 | +write_cdm_csv(filtered, 'our_output.csv') |
| 124 | +print(f'{len(filtered)} conjunctions in {result.elapsed_seconds:.1f}s') |
| 125 | +" |
| 126 | + |
| 127 | +# 4. Compare against answer key: |
| 128 | +# See data/tracss/compare_subset.py (helper script) |
| 129 | +``` |
0 commit comments