Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
gittensor-ai-lab
/
sparkinfer
Public
Notifications
You must be signed in to change notification settings
Fork
79
Star
83
Code
Issues
4
Pull requests
1
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Projects
Security and quality
Insights
Actions: gittensor-ai-lab/sparkinfer
Actions
All workflows
Workflows
build-attested-binaries
build-attested-binaries
cap-open-prs
cap-open-prs
close-stale-prs
close-stale-prs
copycat-guard
copycat-guard
eval-policy
eval-policy
noise-penalty
noise-penalty
pages-build-deployment
pages-build-deployment
publish-docker
publish-docker
rtx5090-required
rtx5090-required
sensitive-paths-guard
sensitive-paths-guard
Show more workflows...
Management
Caches
Deployments
build-attested-binaries
build-attested-binaries
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Show workflow options
Create status badge
Create status badge
Loading
Uh oh!
There was an error while loading.
Please reload this page
.
build-attested-binaries.yml
will be ignored since log searching is not yet available
1,276 workflow runs
1,276 workflow runs
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
perf(bonsai): the decode shadow reads the attention and output projec…
build-attested-binaries
#1332:
Commit
7212971
pushed by
skyrocket2026
In progress
main
main
In progress
View workflow file
perf(kernels): the fused Q4_K prefill GEMM decodes on its own warps a…
build-attested-binaries
#1331:
Commit
2fa6fd5
pushed by
skyrocket2026
12m 53s
main
main
12m 53s
View workflow file
perf(bonsai): the decode shadow reads the attention and output projections in their ternary blocks too, with the GDN norm and the attention gate folded into their rotation (1.11x cb-decode @c16 on Ternary-Bonsai-2)
build-attested-binaries
#1330:
Pull request
#1166
opened by
FranDev132
8m 31s
FranDev132:perf/bonsai-shadow-attn-out
FranDev132:perf/bonsai-shadow-attn-out
8m 31s
View #1166
View workflow file
perf(kernels): the fused Q4_K prefill GEMM decodes on its own warps and takes a 256-row tile past 256 rows (1.28x prefill @512 on Ternary-Bonsai-2)
build-attested-binaries
#1329:
Pull request
#1163
synchronize by
FranDev132
14m 29s
FranDev132:perf/bonsai-prefill-qb-ws-tall
FranDev132:perf/bonsai-prefill-qb-ws-tall
14m 29s
View #1163
View workflow file
perf(bonsai): packed rows read the decode shadow's FFN on the int8 te…
build-attested-binaries
#1328:
Commit
61ccf92
pushed by
skyrocket2026
16m 17s
main
main
16m 17s
View workflow file
perf(kernels): the fused Q4_K prefill GEMM decodes on its own warps and takes a 256-row tile past 256 rows (1.28x prefill @512 on Ternary-Bonsai-2)
build-attested-binaries
#1327:
Pull request
#1163
synchronize by
FranDev132
13m 0s
FranDev132:perf/bonsai-prefill-qb-ws-tall
FranDev132:perf/bonsai-prefill-qb-ws-tall
13m 0s
View #1163
View workflow file
perf(kernels): the fused Q4_K prefill GEMM decodes on its own warps and takes a 256-row tile past 256 rows (1.28x prefill @512 on Ternary-Bonsai-2)
build-attested-binaries
#1326:
Pull request
#1163
synchronize by
FranDev132
10m 3s
FranDev132:perf/bonsai-prefill-qb-ws-tall
FranDev132:perf/bonsai-prefill-qb-ws-tall
10m 3s
View #1163
View workflow file
perf(bonsai): packed rows read the decode shadow's FFN on the int8 tensor cores, one weight read for every row (1.33x cb-decode @c8 on Ternary-Bonsai-2)
build-attested-binaries
#1325:
Pull request
#1164
synchronize by
FranDev132
12m 11s
FranDev132:perf/bonsai-shadow-ternary-tc
FranDev132:perf/bonsai-shadow-ternary-tc
12m 11s
View #1164
View workflow file
perf(bonsai): gate/up stay in their ternary blocks, and packed batche…
build-attested-binaries
#1324:
Commit
dc723a4
pushed by
skyrocket2026
19m 30s
main
main
19m 30s
View workflow file
perf(bonsai): packed rows read the decode shadow's FFN on the int8 tensor cores, one weight read for every row (1.33x cb-decode @c8 on Ternary-Bonsai-2)
build-attested-binaries
#1323:
Pull request
#1164
opened by
FranDev132
12m 24s
FranDev132:perf/bonsai-shadow-ternary-tc
FranDev132:perf/bonsai-shadow-ternary-tc
12m 24s
View #1164
View workflow file
perf(bonsai): gate/up stay in their ternary blocks, and packed batches past 8 rows read the whole FFN there on the int8 tensor cores (1.43x cb-decode @c16 on Ternary-Bonsai-2)
build-attested-binaries
#1322:
Pull request
#1161
synchronize by
FranDev132
11m 37s
FranDev132:perf/bonsai-native-gate-up
FranDev132:perf/bonsai-native-gate-up
11m 37s
View #1161
View workflow file
perf(kernels): the fused Q4_K prefill GEMM decodes on its own warps and takes a 256-row tile past 256 rows (1.28x prefill @512 on Ternary-Bonsai-2)
build-attested-binaries
#1321:
Pull request
#1163
opened by
FranDev132
10m 36s
FranDev132:perf/bonsai-prefill-qb-ws-tall
FranDev132:perf/bonsai-prefill-qb-ws-tall
10m 36s
View #1163
View workflow file
perf(kernels): the long-context int8 prefill GEMM loses its bank conf…
build-attested-binaries
#1320:
Commit
20a4fbd
pushed by
skyrocket2026
13m 13s
main
main
13m 13s
View workflow file
fix(bonsai): release the decode shadow's VRAM for batched prefill, no…
build-attested-binaries
#1319:
Commit
cc6cfc0
pushed by
skyrocket2026
12m 32s
main
main
12m 32s
View workflow file
perf(kernels): the long-context int8 prefill GEMM loses its bank conflicts, takes a 64x64 warp tile and forms the FFN's SwiGLU in its epilogue (1.07x prefill @4k on Ternary-Bonsai-2)
build-attested-binaries
#1318:
Pull request
#1162
opened by
FranDev132
9m 3s
FranDev132:perf/dense-prefill-i8-gemm-w4
FranDev132:perf/dense-prefill-i8-gemm-w4
9m 3s
View #1162
View workflow file
refactor(qwen35_prefill): improve GDN handling and context management…
build-attested-binaries
#1317:
Commit
053ef27
pushed by
skyrocket2026
12m 46s
main
main
12m 46s
View workflow file
fix(bonsai): release the decode shadow's VRAM for batched prefill, not only a new session
build-attested-binaries
#1316:
Pull request
#1160
reopened by
skyrocket2026
12m 34s
inference2026:fix/bonsai-decode-shadow-vram-release
inference2026:fix/bonsai-decode-shadow-vram-release
12m 34s
View #1160
View workflow file
perf(bonsai): gate/up stay in their ternary blocks, and packed batches past 8 rows read the whole FFN there on the int8 tensor cores (1.43x cb-decode @c16 on Ternary-Bonsai-2)
build-attested-binaries
#1315:
Pull request
#1161
opened by
FranDev132
10m 8s
FranDev132:perf/bonsai-native-gate-up
FranDev132:perf/bonsai-native-gate-up
10m 8s
View #1161
View workflow file
fix(bonsai): release the decode shadow's VRAM for batched prefill, not only a new session
build-attested-binaries
#1314:
Pull request
#1160
opened by
inference2026
12m 38s
inference2026:fix/bonsai-decode-shadow-vram-release
inference2026:fix/bonsai-decode-shadow-vram-release
12m 38s
View #1160
View workflow file
perf(kernels): pack bf16 V into the wmma tile for dense-GGUF full attention (1.013x prefill @512 on Ternary-Bonsai-2)
build-attested-binaries
#1313:
Pull request
#1159
opened by
JamesR111
10m 24s
JamesR111:perf/bonsai-prefill-attn-vpack
JamesR111:perf/bonsai-prefill-attn-vpack
10m 24s
View #1159
View workflow file
perf(qwen35): fill the device on the N=512 GDN bf16 GEMMs and fuse th…
build-attested-binaries
#1312:
Commit
c700f98
pushed by
skyrocket2026
13m 56s
main
main
13m 56s
View workflow file
perf(bonsai): decode reads the head and GDN qkv/z from ternary copies through a tensor-core GEMV (1.07x decode @128, 1.08x cb-decode @c4 on Ternary-Bonsai-2)
build-attested-binaries
#1311:
Pull request
#1157
opened by
widecloud
11m 39s
widecloud:perf-bonsai-ternary-decode
widecloud:perf-bonsai-ternary-decode
11m 39s
View #1157
View workflow file
perf(bonsai): decode and small packed batches read the FFN from a ternary copy through a dp4a GEMV (1.35x decode @128)
build-attested-binaries
#1310:
Pull request
#1154
synchronize by
kaivaryn
13m 5s
kaivaryn:perf/bonsai-ffn-shadow-v2
kaivaryn:perf/bonsai-ffn-shadow-v2
13m 5s
View #1154
View workflow file
perf(qwen35): dense-GGUF GDN prefill at 512 takes fused quantized-B, not bf16 (1.29x prefill @512 on Ternary-Bonsai-2)
build-attested-binaries
#1309:
Pull request
#1156
opened by
DripMicro
12m 7s
DripMicro:perf/bonsai-i8-gdn-512
DripMicro:perf/bonsai-i8-gdn-512
12m 7s
View #1156
View workflow file
perf(kernels): dense Q4_K cb-decode takes the tensor-core down and in…
build-attested-binaries
#1308:
Commit
f90d6b5
pushed by
skyrocket2026
12m 56s
main
main
12m 56s
View workflow file
Previous
1
2
3
4
5
…
51
52
Next
You can’t perform that action at this time.