Table of contents for the Iron docs. The top-level README is the curated landing page; this index lists every page so you can jump straight to a topic. New contributors (and their agents) should read Getting started → Developing → Testing before opening a PR.
- Getting started — toolchain, clone, first build, first kernel.
- Developing — repo layout, the
makedev loop, branching, commits, debugging, and the kernel-authoring hazards (⚠️ ) that cause silent or catastrophic failure. Required reading before writing a kernel. - Kernel style guide — the authority on how to write one kernel/bench/test: file shape, naming,
#[kernel(variants(...))], shared primitives, the CPU oracle, the bench. The target style the library is migrating toward. - Testing — the test layers, what runs in CI vs locally, how to write a test, coverage targets, and the gaps in the test infrastructure that let bugs through silently.
- Publishing — the
dev→mainrelease flow.
- Architecture — how a
#[kernel]becomes a compiled shader, and how the bench runner, test runner, and kernel profiling work end-to-end (current in-process vs planned-subprocess execution model). - CLI — the
ironbinary:bench,build,emit,inspect,device,snap,diff. - Kernel audit — per-op coverage table: which MLX / Iron kernels are ported, partial, or still missing, with the gaps and open PRs called out.
Long-form specs and design docs live in specs/.
- Bench metrics spec — planned
iron benchmeasurement additions (latency, GFLOP/s, roofline/utilization, bottleneck) so kernels can actually be optimized and precisions compared; includes the precision-support roadmap (nvfp4/mxfp4/mxfp8) and M5 Neural Accelerator context. - Kernel consolidation plan — the singular roadmap for restructuring
wh-iron-std's kernels: thekernels/<family>/target layout, the three LOC-reduction tools, the provenconv/exemplar, and the family-by-family migration order. (Authoring style lives in the style guide.) - Hy3 MoE prep — kernel-layer readiness for Tencent Hy3 (
hy_v3): router path choice, 192/top-8 shapes, oQ2 int2 expert path, local checkpoint notes. - Toolchain design — the
#[kernel]/#[bench]/#[test_kernel]macro surface and how the IR lowers to MSL. - Proposed optimizations — hot-path patterns that need codegen-layer support, with rationale and implementation sketches.
Codegen-backend specs for taking the DSL beyond Metal. Each is a delta on the CUDA spec, which defines the shared backend seam.
- CUDA / NVIDIA — the backend-seam design (DSL → C++/PTX) the other ports build on.
- AMD / ROCm — HIP/ROCm delta on the CUDA seam.
- Vulkan / SPIR-V — portable compute via SPIR-V.
- Apple ANE — the Neural Engine: why it's not directly programmable and what a port would entail.
- Top-level
README— project landing page. CONTRIBUTING— issue / PR process, agentic-contribution disclosure, code of conduct.