Skip to content

Commit b347e02

Browse files
author
anderdc
committed
serving: attest image split from the runtime; per-card attestation pays N cards per hotkey; sampled audit
- docker/attest is its own image (entrius/gt-attest, attest-image.yml): one long-running container per box beside the runtime, on every GPU; runtime images no longer carry the sidecar. GET /info self-report (telemetry only). - attest.py judges every card the sidecar reports: capacity = passing cards, score = speed credit x cards, UUID dedupe across all cards; a card without the model resident (free VRAM above total - 0.8 x reservation) does not count. - Verification is sampled per (hotkey, round): every baseline prompt and failed request, plus max(10, 20%) of the completed gateway requests. - Loadout carries runtime_image (digest, written by the conformance job) and attest.image; both persisted per round for the /compute release card. SERVING_ATTEST_REFERENCE_URL + compose `attest` service for the reference. - Conformance on Lium rents a second pod for the attest image; check_serving_runtime --attest-url. Claude-Session: https://claude.ai/code/session_01FJtQK6NCjUgYkGTopczEf4
1 parent 6910c52 commit b347e02

20 files changed

Lines changed: 606 additions & 115 deletions

‎.env.example‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -60,5 +60,9 @@ SERVING_ENABLED=false
6060
# Overrides the primary release's reference_url. Unset = reference_url from the loadout, else its audit_bank snapshot.
6161
# SERVING_REFERENCE_URL=http://reference:8080
6262
# SERVING_REFERENCE_API_KEY=<bearer for a remote reference>
63+
# The attest container beside the reference (compose service `attest`, image entrius/gt-attest:GT_ATTEST_TAG).
64+
# Default: the reference host on port 8081 — set it when the sidecar is a separate service or host.
65+
# SERVING_ATTEST_REFERENCE_URL=http://attest:8081
66+
# GT_ATTEST_TAG=v1
6367
# SPARKINFER_TAG=<commit> # image tag for the reference sidecar = runtime_pin commit in the loadout
6468
# SPARKINFER_MODEL_SHA256=<model_sha256 from the loadout release>

‎.github/workflows/attest-image.yml‎

Lines changed: 74 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,74 @@
1+
name: attest image
2+
3+
# Builds the attestation sidecar (docker/attest) and publishes it as entrius/gt-attest:<tag>. One image for every
4+
# blessed release: miners run it beside their runtime container, validators beside their reference. Bump the tag
5+
# (v1, v2, ...) when docker/attest changes and point `attest.image` in serving_loadout.json at it. No GPU in CI: the
6+
# build proves it compiles; the sparkinfer conformance job exercises it on a rented 5090.
7+
8+
on:
9+
workflow_dispatch:
10+
inputs:
11+
image_tag:
12+
description: "Tag to push (e.g. v1). Empty = the tag in attest.image of serving_loadout.json."
13+
required: false
14+
default: ""
15+
cuda_archs:
16+
description: "CMAKE_CUDA_ARCHITECTURES"
17+
required: false
18+
default: "120"
19+
push:
20+
paths:
21+
- docker/attest/**
22+
- .github/workflows/attest-image.yml
23+
24+
jobs:
25+
image:
26+
name: ${{ github.event_name == 'workflow_dispatch' && 'build and push' || 'build only (dry run)' }}
27+
runs-on: ubuntu-latest
28+
steps:
29+
- name: Check out the repo
30+
uses: actions/checkout@v4
31+
32+
- name: Resolve the tag (input, else attest.image in serving_loadout.json)
33+
id: tag
34+
run: |
35+
TAG="${{ inputs.image_tag }}"
36+
if [ -z "$TAG" ]; then
37+
TAG=$(jq -r '.releases[0].attest.image // "entrius/gt-attest:dev"' gittensor/validator/weights/serving_loadout.json | sed 's/.*://')
38+
fi
39+
echo "tag=$TAG" >> "$GITHUB_OUTPUT"
40+
echo "building docker/attest as entrius/gt-attest:$TAG"
41+
42+
- name: Free disk space (CUDA devel image is large)
43+
run: |
44+
sudo rm -rf /usr/share/dotnet /usr/local/lib/android /opt/ghc /opt/hostedtoolcache/CodeQL
45+
docker system prune -af
46+
47+
- name: Log in to Docker Hub
48+
if: github.event_name == 'workflow_dispatch'
49+
uses: docker/login-action@v3
50+
with:
51+
username: ${{ secrets.DOCKER_USERNAME }}
52+
password: ${{ secrets.DOCKER_TOKEN }}
53+
54+
- name: Set up Buildx
55+
uses: docker/setup-buildx-action@v3
56+
57+
- name: Build and push
58+
id: build
59+
uses: docker/build-push-action@v5
60+
with:
61+
context: .
62+
file: docker/attest/Dockerfile
63+
push: ${{ github.event_name == 'workflow_dispatch' }}
64+
build-args: |
65+
CUDA_ARCHS=${{ inputs.cuda_archs || '120' }}
66+
GT_ATTEST_VERSION=${{ steps.tag.outputs.tag }}
67+
tags: |
68+
entrius/gt-attest:${{ steps.tag.outputs.tag }}
69+
cache-from: type=gha
70+
cache-to: type=gha,mode=max
71+
72+
- name: Digest
73+
if: github.event_name == 'workflow_dispatch'
74+
run: echo "entrius/gt-attest:${{ steps.tag.outputs.tag }}@${{ steps.build.outputs.digest }}"

‎.github/workflows/sparkinfer-image.yml‎

Lines changed: 8 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -38,6 +38,7 @@ jobs:
3838
outputs:
3939
ref: ${{ steps.ref.outputs.ref }}
4040
tag: ${{ steps.ref.outputs.tag }}
41+
digest: ${{ steps.build.outputs.digest }}
4142
# push-triggered runs (Dockerfile/workflow changed) only BUILD, to catch Dockerfile breakage in CI;
4243
# the image is pushed to Docker Hub only from the Run workflow button (workflow_dispatch).
4344
name: ${{ github.event_name == 'workflow_dispatch' && 'build and push' || 'build only (dry run)' }}
@@ -73,6 +74,7 @@ jobs:
7374
uses: docker/setup-buildx-action@v3
7475

7576
- name: Build and push
77+
id: build
7678
uses: docker/build-push-action@v5
7779
with:
7880
context: .
@@ -96,6 +98,7 @@ jobs:
9698
LIUM_API_KEY: ${{ secrets.LIUM_API_KEY }}
9799
REF: ${{ needs.image.outputs.ref }}
98100
TAG: ${{ needs.image.outputs.tag }}
101+
DIGEST: ${{ needs.image.outputs.digest }}
99102
steps:
100103
- name: Check out the repo
101104
uses: actions/checkout@v4
@@ -123,11 +126,12 @@ jobs:
123126
path: conformance-${{ env.REF }}/
124127
if-no-files-found: warn
125128

126-
- name: Bump runtime_pin and the release's speed facts
129+
- name: Bump runtime_pin, runtime_image and the release's speed facts
127130
run: |
128-
jq --arg pin "gittensor-ai-lab/sparkinfer@$REF" --arg run "$GITHUB_RUN_ID" \
131+
jq --arg pin "gittensor-ai-lab/sparkinfer@$REF" --arg image "entrius/sparkinfer:$TAG@$DIGEST" --arg run "$GITHUB_RUN_ID" \
129132
--slurpfile speed "conformance-$REF/speed.json" \
130133
'.releases[0].runtime_pin = $pin
134+
| .releases[0].runtime_image = $image
131135
| .releases[0].speed = ((.releases[0].speed // {}) + {
132136
single_stream_decode_tps: $speed[0].single_stream_decode_tps,
133137
decode_per_request: $speed[0].decode_per_request,
@@ -153,7 +157,8 @@ jobs:
153157
body: |
154158
`entrius/sparkinfer:${{ env.REF }}` passed `scripts/check_serving_runtime.py` on a rented RTX 5090
155159
(run ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} — report in the
156-
`serving-conformance-${{ env.REF }}` artifact). This bumps `runtime_pin` and writes the release's
160+
`serving-conformance-${{ env.REF }}` artifact). This bumps `runtime_pin`, pins `runtime_image` to the pushed
161+
image's digest, and writes the release's
157162
blessing-time speed facts (`speed.decode_per_request`, the per-request decode curve served traffic is
158163
priced against) and attestation facts (`attest.*`). Update the sha named in the comments at
159164
`gittensor/constants.py` and `gittensor/serving/audit.py` if the blessing rules changed.

‎docker-compose.vali.yml‎

Lines changed: 17 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -35,7 +35,8 @@ services:
3535
# sparkinfer_server image miners run (built by us from a pinned upstream commit, docker/sparkinfer.Dockerfile),
3636
# on the validator's own GPU, so audits compare against a live honest copy instead of a bank snapshot.
3737
# SPARKINFER_TAG=<runtime_pin commit> docker compose -f docker-compose.vali.yml --profile reference up -d
38-
# then set SERVING_REFERENCE_URL=http://reference:8080 in .env. See the Serving Runtime Contract in the miner docs.
38+
# then set SERVING_REFERENCE_URL=http://reference:8080 and SERVING_ATTEST_REFERENCE_URL=http://attest:8081 in .env.
39+
# See the Serving Runtime Contract in the miner docs.
3940
reference:
4041
image: entrius/sparkinfer:${SPARKINFER_TAG:-latest}
4142
container_name: gt-vali-reference
@@ -56,6 +57,21 @@ services:
5657
volumes:
5758
- ./data/models:/opt/sparkinfer/models
5859

60+
# Attestation sidecar beside the reference (profile "reference"): the validator asks it for the expected digest
61+
# and reference wall time of each round's challenge. Same image every miner box runs (attest.image in the loadout).
62+
attest:
63+
image: entrius/gt-attest:${GT_ATTEST_TAG:-v1}
64+
container_name: gt-vali-attest
65+
restart: unless-stopped
66+
profiles: ["reference"]
67+
deploy:
68+
resources:
69+
reservations:
70+
devices:
71+
- driver: nvidia
72+
count: 1
73+
capabilities: [gpu]
74+
5975
watchtower:
6076
image: nickfedor/watchtower
6177
container_name: gittensor-watchtower

‎docker/attest/Dockerfile‎

Lines changed: 39 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,39 @@
1+
# syntax=docker/dockerfile:1.7
2+
# Gittensor attestation sidecar: one long-running container per miner box (and beside a validator's reference), on
3+
# every GPU the box has. Answers the validator's hardware challenge (gt_attest: seeded VRAM fill + deterministic
4+
# GEMM chain, docker/attest/gt_attest.cu) and a self-report (GET /info). Runtime-agnostic — the same image serves
5+
# every blessed release, so blessing a new runtime never touches it. Pinned by `attest.image` in the serving loadout.
6+
#
7+
# docker build -f docker/attest/Dockerfile -t entrius/gt-attest:<tag> .
8+
# docker run -d --name gt-attest --gpus all -p 8081:8081 entrius/gt-attest:<tag>
9+
10+
ARG CUDA_VERSION=12.8.1
11+
ARG UBUNTU_VERSION=24.04
12+
13+
FROM nvidia/cuda:${CUDA_VERSION}-devel-ubuntu${UBUNTU_VERSION} AS build
14+
# sm_120 = RTX 5090. The digest is only comparable across cards of one architecture; add more only if the fleet
15+
# blesses other hardware.
16+
ARG CUDA_ARCHS=120
17+
COPY docker/attest/gt_attest.cu /src/gt_attest.cu
18+
# -fmad=false: the summation is bit-identical on every card of the architecture (the validator's reference recomputes
19+
# the digest).
20+
RUN nvcc -O3 -fmad=false -gencode arch=compute_${CUDA_ARCHS},code=sm_${CUDA_ARCHS} \
21+
-o /src/gt_attest /src/gt_attest.cu -lnvidia-ml
22+
23+
FROM nvidia/cuda:${CUDA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}
24+
ENV DEBIAN_FRONTEND=noninteractive
25+
RUN apt-get update && apt-get install -y --no-install-recommends python3 curl \
26+
&& rm -rf /var/lib/apt/lists/*
27+
WORKDIR /opt/gt-attest
28+
COPY --from=build /src/gt_attest bin/gt_attest
29+
COPY docker/attest/attest_server.py bin/attest_server.py
30+
ARG GT_ATTEST_VERSION=dev
31+
ENV GT_ATTEST_BIN=/opt/gt-attest/bin/gt_attest \
32+
GT_ATTEST_VERSION=${GT_ATTEST_VERSION} \
33+
ATTEST_PORT=8081
34+
EXPOSE 8081
35+
LABEL io.gittensor.serving.attest="gt-attest" \
36+
io.gittensor.serving.attest_version="${GT_ATTEST_VERSION}"
37+
HEALTHCHECK --interval=30s --timeout=5s --start-period=30s --retries=3 \
38+
CMD curl -fsS http://127.0.0.1:8081/info || exit 1
39+
ENTRYPOINT ["python3", "bin/attest_server.py"]

‎docker/attest/attest_server.py‎

Lines changed: 62 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,14 @@
11
#!/usr/bin/env python3
2-
"""Attestation sidecar for the serving runtime image: runs gt_attest on request (stdlib only).
2+
"""Attestation sidecar (image entrius/gt-attest, docker/attest/Dockerfile): runs gt_attest on request (stdlib only).
33
44
POST /v1/attest {"seed": <u64>, "iters": <n>, "fill": true} -> gt_attest's JSON plus "queued_ms"
55
GET /v1/attest/info -> devices without running a challenge (iters 1, no fill,
66
small working set)
7+
GET /info -> self-report: GPUs as nvidia-smi sees them, driver,
8+
host CPU/RAM, sidecar version. Telemetry for the
9+
operator and the /compute page — a host can say
10+
anything here, so nothing is paid on it; the
11+
challenge above is the proof.
712
Challenges run one at a time; concurrent requests queue and report their own wall including the wait, which is what
813
lets a validator see two hotkeys sharing one card. Optional bearer via ATTEST_API_KEY.
914
"""
@@ -15,7 +20,8 @@
1520
import time
1621
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
1722

18-
BIN = os.environ.get('GT_ATTEST_BIN', '/opt/sparkinfer/bin/gt_attest')
23+
BIN = os.environ.get('GT_ATTEST_BIN', '/opt/gt-attest/bin/gt_attest')
24+
VERSION = os.environ.get('GT_ATTEST_VERSION', 'dev')
1925
PORT = int(os.environ.get('ATTEST_PORT', '8081'))
2026
API_KEY = os.environ.get('ATTEST_API_KEY')
2127
LOCK = threading.Lock()
@@ -37,6 +43,55 @@ def run(args, timeout=120.0):
3743
return 200, out
3844

3945

46+
def self_report() -> dict:
47+
query = 'uuid,name,driver_version,memory.total,memory.used,pci.bus_id,temperature.gpu,utilization.gpu'
48+
gpus, error = [], None
49+
try:
50+
out = subprocess.run(
51+
['nvidia-smi', f'--query-gpu={query}', '--format=csv,noheader,nounits'],
52+
capture_output=True,
53+
text=True,
54+
timeout=10.0,
55+
)
56+
if out.returncode != 0:
57+
error = (out.stderr or out.stdout).strip()[:300]
58+
for line in out.stdout.splitlines():
59+
cols = [c.strip() for c in line.split(',')]
60+
if len(cols) == 8:
61+
gpus.append(
62+
{
63+
'uuid': cols[0],
64+
'name': cols[1],
65+
'driver': cols[2],
66+
'memory_total_mib': int(cols[3]),
67+
'memory_used_mib': int(cols[4]),
68+
'pci_bus_id': cols[5],
69+
'temperature_c': int(cols[6]),
70+
'utilization_pct': int(cols[7]),
71+
}
72+
)
73+
except Exception as e: # no nvidia-smi in the container, or a hung driver
74+
error = repr(e)[:300]
75+
mem_total_kib = None
76+
try:
77+
with open('/proc/meminfo') as f:
78+
for line in f:
79+
if line.startswith('MemTotal:'):
80+
mem_total_kib = int(line.split()[1])
81+
break
82+
except OSError:
83+
pass
84+
return {
85+
'version': VERSION,
86+
'hostname': os.uname().nodename,
87+
'cpus': os.cpu_count(),
88+
'mem_total_gib': round(mem_total_kib / 1048576, 1) if mem_total_kib else None,
89+
'gpus': gpus,
90+
'error': error,
91+
'ts': time.time(),
92+
}
93+
94+
4095
class Handler(BaseHTTPRequestHandler):
4196
def _send(self, status, payload):
4297
body = json.dumps(payload).encode()
@@ -53,7 +108,11 @@ def _auth(self):
53108
return True
54109

55110
def do_GET(self):
56-
if self.path != '/v1/attest/info' or not self._auth():
111+
if not self._auth():
112+
return
113+
if self.path == '/info':
114+
return self._send(200, self_report())
115+
if self.path != '/v1/attest/info':
57116
return self._send(404, {'error': 'not found'})
58117
self._send(
59118
*run(['--seed', '1', '--iters', '1', '--dim', '256', '--matrices', '2', '--device', 'all'], timeout=30.0)

‎docker/attest/entrypoint.sh‎

Lines changed: 0 additions & 4 deletions
This file was deleted.

‎docker/sparkinfer.Dockerfile‎

Lines changed: 6 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,7 @@
99
#
1010
# docker build -f docker/sparkinfer.Dockerfile --build-arg SPARKINFER_REF=<commit> -t entrius/sparkinfer:<commit> .
1111
# docker run --gpus all -p 8080:8080 -v $PWD/data/models:/opt/sparkinfer/models entrius/sparkinfer:<commit>
12+
# A miner box also runs the attestation container (docker/attest, entrius/gt-attest) beside it.
1213

1314
ARG CUDA_VERSION=12.8.1
1415
ARG UBUNTU_VERSION=24.04
@@ -38,11 +39,6 @@ RUN git clone ${SPARKINFER_REPO} sparkinfer \
3839
RUN cmake -S sparkinfer -B sparkinfer/build -G Ninja \
3940
-DCMAKE_BUILD_TYPE=Release -DBUILD_SERVER=ON -DCMAKE_CUDA_ARCHITECTURES=${CUDA_ARCHS} \
4041
&& cmake --build sparkinfer/build
41-
# Hardware attestation challenge (docker/attest): deterministic fp32 GEMM chain + VRAM fill. -fmad=false so the
42-
# summation is bit-identical on every card of the architecture (the validator's reference recomputes the digest).
43-
COPY docker/attest/gt_attest.cu /src/attest/gt_attest.cu
44-
RUN nvcc -O3 -fmad=false -gencode arch=compute_${CUDA_ARCHS},code=sm_${CUDA_ARCHS} \
45-
-o /src/attest/gt_attest /src/attest/gt_attest.cu -lnvidia-ml
4642

4743
# ---------- runtime ----------
4844
FROM nvidia/cuda:${CUDA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}
@@ -60,9 +56,6 @@ COPY --from=build /src/sparkinfer/build/runtime build/runtime
6056
COPY --from=build /src/sparkinfer/build/moe build/moe
6157
COPY --from=build /src/sparkinfer/build/server build/server
6258
COPY --from=build /src/SPARKINFER_COMMIT SPARKINFER_COMMIT
63-
COPY --from=build /src/attest/gt_attest bin/gt_attest
64-
COPY docker/attest/attest_server.py bin/attest_server.py
65-
COPY docker/attest/entrypoint.sh bin/entrypoint.sh
6659
# sparkinfer_server links libsparkinfer_runtime.so + libsparkinfer_moe.so (ldd-verified on 1b8b962, 9e43bfa).
6760
ENV LD_LIBRARY_PATH=/opt/sparkinfer/build/runtime:/opt/sparkinfer/build/moe:${LD_LIBRARY_PATH}
6861

@@ -81,15 +74,15 @@ ENV SPARKINFER_REF=${SPARKINFER_REF} \
8174
HOST=0.0.0.0 \
8275
PORT=8080
8376
VOLUME ["/opt/sparkinfer/models"]
84-
# 8080 inference, 8081 attestation sidecar (docker/attest/attest_server.py)
85-
EXPOSE 8080 8081
77+
# 8080 inference. Hardware attestation is the separate entrius/gt-attest container (docker/attest), not this image.
78+
EXPOSE 8080
8679
LABEL org.opencontainers.image.source="https://github.com/gittensor-ai-lab/sparkinfer" \
8780
io.gittensor.serving.runtime="sparkinfer" \
8881
io.gittensor.serving.runtime_ref="${SPARKINFER_REF}"
8982

9083
HEALTHCHECK --interval=30s --timeout=5s --start-period=600s --retries=5 \
9184
CMD curl -fsS http://127.0.0.1:8080/v1/models || exit 1
9285

93-
# bin/entrypoint.sh starts the attestation sidecar (:8081), then run.sh downloads the blessed model + tokenizer into
94-
# MODELS_DIR on first start (--download) and execs the prebuilt server binary.
95-
ENTRYPOINT ["bash", "bin/entrypoint.sh"]
86+
# run.sh downloads the blessed model + tokenizer into MODELS_DIR on first start (--download) and execs the prebuilt
87+
# server binary.
88+
ENTRYPOINT ["bash", "server/run.sh", "--download"]

‎gittensor/constants.py‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -188,6 +188,10 @@
188188
SERVING_ATTEST_MIN_FILL_RATIO = 0.6
189189
SERVING_ATTEST_TIMEOUT = 45.0 # seconds: fill + chain on a 5090 is ~1.5 s; queued challenges show up as slow
190190
SERVING_ATTEST_UUID_MEMORY_ROUNDS = 12
191+
# A card counts only with the model resident: free VRAM before the fill must be at most total minus this fraction of
192+
# the reservation (a bare 5090 shows ~32 GB free, one holding the model ~8 GB). Every passing card is a card-hour, so
193+
# a multi-GPU hotkey earns per card it actually serves from — a second card with nothing loaded earns nothing.
194+
SERVING_ATTEST_MODEL_RESIDENT_RATIO = 0.8
191195
SERVING_VRAM_MODEL_RESERVED_BYTES = 24e9 # what sparkinfer holds with the model loaded (7498736 on a 5090: 23.7 GB)
192196
SERVING_VERIFY_WORKERS = 8 # concurrent /v1/score calls to the reference per round
193197
# Latency credit on the validator-observed time to first streamed token (network + queue + prefill). An honest
@@ -231,6 +235,12 @@
231235
# python scripts/serving_cheat_experiment.py analyze --honest <honest rows> --cheaters <cheater rows>
232236
SERVING_AUDIT_WINDOW = 10
233237
SERVING_AUDIT_WINDOW_THRESHOLDS = ((1, 0.8),)
238+
# Verification is one reference prefill per request, so it is sampled: per (hotkey, round) every baseline prompt and
239+
# every failed request is judged, and of the completed gateway requests a random max(SAMPLE_MIN, SAMPLE_FRACTION x n)
240+
# — the floor fills a window in one round, the fraction bounds reference load under real traffic. Nothing on the
241+
# wire says which requests were drawn.
242+
SERVING_AUDIT_SAMPLE_FRACTION = 0.2
243+
SERVING_AUDIT_SAMPLE_MIN = 10
234244
SERVING_QUARANTINE_S = 3600.0
235245
# Baseline traffic: every round the validator sends each serving axon this many baseline prompts of its own, at
236246
# random moments spread over the round (so they do not mark the round boundary), over the same path as user

0 commit comments

Comments
 (0)