docs(website): make the self-improvement page describe only what ships #689
Workflow file for this run
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # Static analysis (CodeQL) — the security-scanning half of the supply-chain | |
| # posture AGENTS.md describes: `cargo deny` gates advisories/bans/sources/ | |
| # licenses, `dependency-review.yml` reviews every PR's dependency diff, and | |
| # this is what actually reads the code. | |
| # | |
| # ── Why this file exists at all (#3260) ────────────────────────────────────── | |
| # | |
| # Until this file, CodeQL ran as GitHub's managed "default setup" — no | |
| # committed workflow, configured entirely through repository settings. Its | |
| # last successful analysis was 2026-07-13; every run after that, on | |
| # 2026-07-17, failed in five to nine seconds each. That is too fast to be a | |
| # Rust autobuild actually attempting this workspace — a real build failure | |
| # would cost minutes, not single digits of seconds — which points at an | |
| # init-time configuration error rather than a build problem, though the exact | |
| # cause is not recoverable: default setup's run logs are not retained the way | |
| # a committed workflow's are, and by the time #3260 was filed they had already | |
| # expired. What is verifiable is that default setup's own state is | |
| # `not-configured` today (`gh api repos/<repo>/code-scanning/default-setup`) — | |
| # GitHub disables default setup automatically after enough consecutive | |
| # failures — so from some point after 2026-07-17 there were not even any red | |
| # runs left to notice. The scanner was not merely failing; for over a month, | |
| # nothing invoked it at all, and nothing in this repository's own gate reads | |
| # CodeQL's status, so nothing said so. | |
| # | |
| # Committing this file switches CodeQL to "advanced setup": a workflow this | |
| # repository owns, whose logs live in this repository's own Actions history | |
| # with the same retention as every other workflow here, and whose language | |
| # list and build step are declared rather than configured through a settings | |
| # page nobody was watching. | |
| # | |
| # ── Why these four languages, and not default setup's seven ───────────────── | |
| # | |
| # Default setup's configured language list (readable via the same API call | |
| # above) was `[actions, javascript, javascript-typescript, python, ruby, rust, | |
| # typescript]` — `javascript`, `javascript-typescript` and `typescript` all | |
| # three at once. Current CodeQL does not accept that combination (the | |
| # consolidated `javascript-typescript` id supersedes the separate | |
| # `javascript`/`typescript` ones), and duplicated/overlapping language ids is | |
| # exactly the kind of error that would fail at init time, before any build | |
| # runs — consistent with the five-to-nine-second failures above, though this | |
| # is inference rather than something a log line confirms. | |
| # | |
| # The four kept here are the languages this repository actually has real | |
| # source in: `rust` (the workspace itself), `python` (`bench/`, `scripts/`), | |
| # `javascript-typescript` (`website/`) and `actions` (every workflow under | |
| # `.github/workflows/`). `ruby` is dropped — the repository's only `.rb` file | |
| # is `packaging/homebrew/stella.rb`, not worth its own CodeQL extractor pass. | |
| # | |
| # ── Why every language here gets `build-mode: none` ───────────────────────── | |
| # | |
| # CodeQL's Rust extractor accepts one build mode, `none`: it extracts from | |
| # source with the workspace resolved by cargo, rather than tracing a compiler | |
| # the way the C/C++/Java extractors do. `manual` is refused outright by | |
| # `codeql database init` — "Rust does not support the manual build mode. | |
| # Please try using one of the following build modes instead: none" — so the | |
| # rust leg carries no build step. `python` and `javascript-typescript` are | |
| # interpreted and `actions` parses YAML directly, so none of those needs one | |
| # either. | |
| # | |
| # The rust leg still installs the toolchain `rust-toolchain.toml` pins, | |
| # because the extractor resolves the workspace through cargo before it | |
| # extracts; a runner with no toolchain hands it an unresolved dependency | |
| # graph. Whether this workspace *compiles* is `ci.yml`'s question, not this | |
| # file's. | |
| # | |
| # ── Why this stays advisory (#3260 DoD item 4) ─────────────────────────────── | |
| # | |
| # Recorded on the issue: advisory, not a required check. This repository | |
| # already gates hard on `cargo deny` as a required check, and adding a second | |
| # required check before a CodeQL run is proven reliably green risks | |
| # reproducing the exact failure this file exists to fix — a scanner nobody can | |
| # get green, blocking every PR, is worse than one running silently. The | |
| # `report` job below is the other half of that decision: the silence itself | |
| # was the defect (five and a half weeks with nothing surfacing it), so a red | |
| # run on `main` now opens a labelled tracking issue and a green run closes it | |
| # — `scripts/codeql-canary.sh`, modeled on `main-canary.yml`'s pattern but | |
| # scoped to the one fact this workflow already knows: did the analyze job | |
| # conclude success or failure. | |
| # | |
| # ── What `rust/cleartext-logging` gets wrong here (#6388) ──────────────────── | |
| # | |
| # That rule had 52 open alerts when #6388 was filed, all high severity, none | |
| # dismissed — a signal firing often enough that nobody read it, so the next | |
| # real one would land in the pile and look the same. Fifty were open by the | |
| # time the sweep ran, because the set moves: three had been answered on | |
| # #6384, a few had closed themselves when the code they pointed at moved, and | |
| # a moved line mints a new alert. Every one of the fifty was read against the | |
| # code, and every one was wrong, in three shapes. They are written down here | |
| # so the next author dismisses by citing this and does not read the same | |
| # fifty sites again. | |
| # | |
| # The rule keys on the NAME of the taint source, not on what the value is, so | |
| # it is wrong three ways: | |
| # | |
| # 1. A correlation id read as a session token. `ZaiProvider::session_id`, | |
| # `lane_journal_key`, the self-driving loop's `sd-<secs>-<pid>` and the | |
| # driver's `drive-<secs>-<hex>` all read as credentials. None grants | |
| # access. `ZaiProvider`'s goes on the wire in the request body under the | |
| # `openrouter` identity, where it is the sticky-routing key | |
| # (`build_body`'s `self.serves_openrouter().then_some(...)`), so hiding | |
| # it from a log protects nothing. | |
| # 2. A guard read as a leak. `validate_project_secrets`, | |
| # `refuse_unless_trusted`, `no_api_key_error` and `api_key_refusal` are | |
| # the functions that turn a credential DOWN. What they return is a | |
| # refusal naming a flag, a path or an environment variable's name. The | |
| # tests that pin those refusals assert the key is absent, and the string | |
| # CodeQL flags is the assertion message that prints only when the assert | |
| # fails. | |
| # 3. A keyboard key read as a cryptographic one. Every `stella-tui` alert | |
| # traces to `handle_deck_key`, `handle_session_key` or | |
| # `handle_mcp_auth_key`, whose `key: KeyEvent` is a keystroke. The one | |
| # credential those handlers carry — what a user types into the MCP auth | |
| # prompt — is wrapped in `Secret`, whose `Debug` prints `<redacted>` | |
| # (`crates/stella-tui/src/envelope/workspace_input.rs`) and whose `Drop` | |
| # zeroizes (`AuthPrompt` in `crates/stella-tui/src/views/mcp_tab.rs`). | |
| # `deck_ui/tests/queue.rs` asserts a typed key never appears under | |
| # `Debug`, and CodeQL flagged the panic arm of that very test. | |
| # | |
| # No query filter was added. CodeQL filters by query id, not by path, so the | |
| # only filter that clears these is `exclude: id: rust/cleartext-logging`, | |
| # which would hide the first true one too. A dismissal is per-alert and | |
| # carries a written reason, which is what a later reader needs. | |
| # | |
| # The advisory posture above is unchanged, for a reason #3260 did not have. | |
| # The backlog is empty now, so a required check would cost nothing today. | |
| # What argues against it is the hit rate: of the alerts this query has ever | |
| # produced here, the fifty a person read were wrong, and the rest closed on | |
| # their own when the code they pointed at moved. Make CodeQL required once a | |
| # run produces an alert somebody agrees with. | |
| name: CodeQL | |
| on: | |
| push: | |
| branches: [main] | |
| pull_request: | |
| branches: [main] | |
| schedule: | |
| # Weekly, offset from scorecard.yml's Monday 06:00 so the two scheduled | |
| # security scans do not queue against each other. | |
| - cron: "17 4 * * 1" | |
| workflow_dispatch: | |
| permissions: | |
| contents: read | |
| jobs: | |
| analyze: | |
| name: Analyze (${{ matrix.language }}) | |
| runs-on: ubuntu-latest | |
| timeout-minutes: 60 | |
| permissions: | |
| contents: read | |
| security-events: write # upload SARIF to code scanning | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| include: | |
| - language: rust | |
| build-mode: none | |
| - language: python | |
| build-mode: none | |
| - language: javascript-typescript | |
| build-mode: none | |
| - language: actions | |
| build-mode: none | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7 | |
| # Only the rust leg needs a toolchain — CodeQL's extractor resolves the | |
| # workspace through cargo. The pin comes from rust-toolchain.toml. | |
| - name: Install Rust (rustup; the pin comes from rust-toolchain.toml) | |
| if: matrix.language == 'rust' | |
| uses: dtolnay/rust-toolchain@4cda84d5c5c54efe2404f9d843567869ab1699d4 # stable | |
| - if: matrix.language == 'rust' | |
| uses: Swatinem/rust-cache@f0d9c3887740aee45f6153b24b3a6b815192ec16 # v2 | |
| - name: Initialize CodeQL | |
| uses: github/codeql-action/init@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4 | |
| with: | |
| languages: ${{ matrix.language }} | |
| build-mode: ${{ matrix.build-mode }} | |
| - name: Perform CodeQL analysis | |
| uses: github/codeql-action/analyze@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4 | |
| with: | |
| category: "/language:${{ matrix.language }}" | |
| # Makes a sustained red visible (#3260 DoD item 4) instead of another month | |
| # of silence. Runs only outside `pull_request`: a PR's `GITHUB_TOKEN` cannot | |
| # be trusted with `issues: write` the way a push-to-main or scheduled run's | |
| # can, and per-PR issue noise is not what this is for — the canary is | |
| # answering "is CodeQL still working on `main`", not judging any one PR's | |
| # code. | |
| report: | |
| name: report CodeQL status | |
| needs: analyze | |
| if: ${{ !cancelled() && github.event_name != 'pull_request' }} | |
| runs-on: ubuntu-latest | |
| timeout-minutes: 5 | |
| permissions: | |
| contents: read | |
| issues: write | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7 | |
| - name: Open, refresh or close the codeql-red tracking issue | |
| env: | |
| GH_TOKEN: ${{ github.token }} | |
| run: | | |
| conclusion=failure | |
| if [ "${{ needs.analyze.result }}" = "success" ]; then | |
| conclusion=success | |
| fi | |
| ./scripts/codeql-canary.sh \ | |
| --announce \ | |
| --conclusion "$conclusion" \ | |
| --sha "${{ github.sha }}" \ | |
| --run-url "${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}" |