Skip to content

Latest commit

 

History

History
1182 lines (952 loc) · 360 KB

File metadata and controls

1182 lines (952 loc) · 360 KB

Changelog

All notable changes to the YUP project will be documented in this file.

The format is based on Keep a Changelog.


[2.0.0] - Unreleased

  • MidiKeyboardComponent: multitouch (each finger holds, slides and releases its own note), [<] / [>] octave scroll buttons replacing the - / + selector, a visible range (setVisibleRange, setHighestVisibleKey, onVisibleRangeChanged) inside the available range, and the JUCE velocity sensitivity, channel mask, orientation, black note proportions, key hooks and note label APIs. Behavior change: typed notes follow the new setKeyPressBaseOctave (setOctaveForMiddleC only names octaves), wheel scroll and zoom stay inside the available range, position velocity is scaled by setVelocity and correct for vertical keyboards, a default keyboard shows the scroll buttons, focusLost keeps mouse and touch notes held, vertical keyboards anchor black keys at the back edge, and the back-edge shadow and front-edge line now paint over the white keys instead of under them.

  • ComponentNative::setVsyncEnabled() / isVsyncEnabled() switch vsync at runtime (GL swap interval, falling back from adaptive to plain vsync when the driver refuses it, Metal, D3D; Dawn keeps its creation-time mode). On desktop, vsync on runs at the display rate and vsync off paces to the desired frame rate. On the web, vsync on renders on every display refresh and vsync off paces to the desired frame rate on elapsed time, so 60fps on a 144Hz display averages 60 instead of snapping to 72. The Emscripten page shell gains a VSync toggle and an FPS field (greyed and showing the measured rate while vsync is on)

  • Drawable: a <use> without its own fill or stroke now draws the referenced shape with that shape's gradient, pattern, opacity, filter, clip and markers instead of a flat copy of its color. <use> of <symbol> or <defs> content now renders. On the rive MSAA backend (WebGL2), a draw faded with Graphics::setOpacity is no longer treated as opaque, so solid colors and opaque gradients blend instead of replacing what is below (any component, not only SVG).

  • GpuComputePipeline::compileFromBundle resolves the main0 kernel SPIRV-Cross emits for a GLSL main on Metal, like GpuPipeline::compileFromBundle already did, instead of failing with "Metal compute function not found: main".

  • CMake: yup_add_shader_bundle accepts a COMPUTE stage, and with BUNDLE_RESOURCE ships the .ysl through BUNDLE_RESOURCES instead of embedding it. Bundles are only regenerated when the tool, the arguments or an input (stages and the new DEPENDS) changed. The graphics example declares the modules, data and precompiled shaders of each demo, so a single-demo build (YUP_EXAMPLE_GRAPHICS_DEMO) only links what that demo uses. Only the SpinningCube and GpuAudio demos, which edit shaders live, still need the shader transpiler.

  • Behavior change SystemStats: isOperatingSystem64Bit reports the OS instead of the build (32-bit builds on 64-bit Linux, Android, Windows on ARM; iOS now true; best-effort host probe on WebAssembly). Added MacOS_15, MacOS_26 and MacOS_27. The macOS name reads macOS <version>, the Android version is Build.VERSION.RELEASE and the device description drops the serial, the Linux version is the kernel release, and WebBrowser no longer aliases the MacOSX bit (compare it with ==, or test the family with & WASM).

  • Behavior change Graphics::setClipPath: the clip is now in local coordinates like every draw call, including the drawing area offset, and getClipPath returns it in the current local space. Code that mapped the clip by getTransform().translated (getDrawingArea().getTopLeft()) or by top-level bounds must drop that compensation. Clips on components with an offset parent (for example Drawable <image> clips) now land where they are drawn.

  • Component: a rotated or sheared component, and everything inside it, is now clipped to its real outline instead of the bounding box of its transformed bounds, so children no longer spill past the corners of a skewed parent.

  • SyncSpectralResampler: the per-harmonic accumulation goes through FloatVectorOperations again instead of a hand-written SIMDRegister loop, which was many times slower in debug builds. The graphics synthesizer example caps a per-voice synced series at the note's Nyquist harmonic count.

  • Graphics synthesizer example: PRISM logo and larger buttons in the header, LFO / scope columns aligned with filter / envelopes, and pitch bend and mod wheels beside the keyboard. The mod wheel is a new MOD WHEEL source in the modulation matrix. The matrix now has 16 slots.

  • CMake: retry failed upstream module and validation tool downloads, verify sha256 after download, and fail at configure time when an upstream archive extracts nothing instead of later with missing headers.

  • YDSP backend (emscripten): mint kernel handles from a module-wide counter instead of a JS-realm-local one, so a realm that runs a graph can no longer find another realm's kernel under the same key and silently invoke the wrong module.

  • YDSP VS Code extension: audition the active patch through yup_dsp_compiler run from a Patch Player sidebar view (transport, workspace patch list, audio/MIDI device selects, sample rate, block size, test note), with a single pinned player per window, a status bar, a dedicated playback output channel and an opt-in follow-active-patch mode.

  • YDSP player: --hotreload now works for a standalone .ydsp as well as a .ydsp-project, driven by YdspDiagnostics::getSourceIds(), which reports the source closure of the last compile (root source, project sources and transitively imported files).

  • YDSP diagnostics: check unused library and processor function bodies during editor validation, including return expressions.

  • YDSP tooling: validate standalone processor/function libraries without a graph; add player error details, verbose device/activity reporting, and a test-note option for diagnosing silent playback.

  • YDSP tooling: compile project bundles, validate unsaved project imports in VS Code, and add audio/MIDI device selection and project hot reload to the command-line player.

  • YDSP: add YAML .ydsp-project manifests with patch metadata, explicit-import source inventories, and overridable processor or graph entry points.

  • YDSP: keep the source value of a float literal that adapts to a float64 context, instead of rounding the constant through float32 first; regenerate bundles for codegen revision 14.

  • YDSP: preserve qualified function calls when nesting library imports, including calls from processors and other library functions.

  • YDSP: preserve source ranges through diagnostics and imports, print path:line:column with five-line context and caret underlines, and improve malformed-number and parser errors.

  • YDSP: add integer/boolean match statements with scoped arms, single selector evaluation, optional _ fallback, constant-arm elimination, and invalid-pattern diagnostics; regenerate bundles for language version 4/codegen revision 13.

  • YDSP: add opt-in trace("value={x}") with bounded allocation-free recording, off-thread formatting/console printing, and native/WebAssembly bundle support; regenerate bundles for ABI 2/codegen revision 12.

  • YDSP examples: restore wrapping noise sequences in DigitalDrums, WaveLab and ControlRateWah after the integer saturation change.

  • YDSP: saturate integer kernel add/subtract/multiply across native, WebAssembly and constant folding while retaining direct arithmetic for proven-safe operations; regenerate bundles for codegen revision 11.

  • YDSP: eliminate array bounds checks for proven nonoverflowing integer products and strided indices; regenerate bundles for codegen revision 10.

  • YDSP: preserve integer storage width during constant folding so arithmetic and subsequent comparisons agree with runtime execution; regenerate bundles for codegen revision 9.

  • YDSP: saturate float-to-integer conversions and map NaN to zero across native/WebAssembly and constant folding, preserving direct instructions for proven-safe casts; regenerate bundles for codegen revision 8.

  • YDSP: respect source and destination widths in constant-folded numeric conversions and leave exceptional float-to-int casts unfolded; regenerate bundles for codegen revision 7.

  • YDSP: make constant-folded integer shifts match native/WebAssembly operand widths and masked counts without adding runtime checks; regenerate bundles for codegen revision 6.

  • YDSP: remove proven ring-counter and wrapped delay-tap bounds checks after auditing initialization and event writes; regenerate bundles for codegen revision 5.

  • YDSP: remove redundant bounds checks proven by scalar integer masks, clamps, min/max, selects and nonoverflowing arithmetic; retain checks for mutable or unknown ranges.

  • YDSP: fuse shared comparisons into native selects when every consumer preserves the operands, avoiding temporary bounds booleans.

  • YDSP: use safe-index selection for branchless checked reads from nonempty fixed state arrays.

  • YDSP: compact array bounds checks to unsigned comparisons and hoist delay clamps preceding guarded accesses.

  • YDSP: allow unconditional invariant hoisting and unrelated branch optimization in guarded kernels, preserving short-circuit and bounds checks; enable benchmark ASM dumps with YUP_DUMP_KERNEL=1.

  • YDSP: fold constant int32-to-int64 sign extension so widened constants participate in integer arithmetic folding.

  • YDSP: saturate integer kernel negation, absolute value and overflowing division at both widths in native/WebAssembly code and constant folding; regenerate bundles for codegen revision 4.

  • YDSP: preserve exact signed 64-bit literals, reject overflowing structural sizes and explicit nonfinite constants; regenerate bundles for codegen revision 3.

  • YDSP: preserve vectorization for proven in-range blockSize - N stream loops; retain guards when fixed loop and buffer lengths can differ.

  • YDSP: guard dynamic state-array, struct-field and indexed-stream accesses, diagnose constant invalid indices, and reject overflowing declared state layouts before flattening.

  • YDSP: lower logical operators and ternaries with short-circuit branches, preserving eager select() and reporting conservative optimization restrictions.

  • YDSP: replace input value / output value with input parameter / output parameter in processors and graphs, freeing value as an identifier, with migration diagnostics and updated examples/editor support; regenerate language-version-2 bundles.

  • YDSP: fix kernel listing histograms misreading hex-like mnemonics as machine-code bytes; exclude labels and assembler directives.

  • YDSP: retain cached state-array loads across provably disjoint stores using the same index, with type, overlap, and index-redefinition coverage.

  • YDSP: strengthen scalar tanh accuracy and feedback regression coverage; retain libm after rejecting a bounded ARM64 approximation that slowed the tanh shaper.

  • YDSP: make array store-to-load forwarding type- and lane-aware, recognize disjoint constant-index ranges, and run this cleanup after vectorization/unrolling with overlap regression coverage.

  • YDSP: allow bounded two-iteration unrolling around four-lane native math calls with a conservative call-liveness budget; add output/state parity and rejection coverage.

  • YDSP: expand default-policy timing to all 19 patches and four block sizes, separating prepared processing from reset/setup; add median/spread reporting, run metadata, and previous-JIT log comparison.

  • YDSP: replace the incomplete Zita example with the full steady-state stereo network and EQs; add a local JIT benchmark with reference parameter settings, native listings, impulse dumps, and impulse-tail regression coverage.

  • YDSP: hoist and share vector math call targets across native kernels while preserving strict/fastMath accuracy selection.

  • YDSP: omit unused ARM64 array-base setup and reuse dead local vector FMA addends on ARM64/x64.

  • YDSP: preserve unique temporaries and specialize indices in unrolled loops, enabling subsequent FMA contraction and constant vector array offsets on ARM64/x64.

  • YDSP: reuse state-array reads between writes, encode constant scalar array offsets on ARM64/x64, combine constant multiplication chains under fastMath, and rematerialize dry/wet coefficients after math calls.

  • YDSP: vectorize if-converted stream loops and positive constant starts, preserve fixed-index array dependencies, and hoist iteration-local invariant clamp/conversion chains.

  • YDSP: extend fastMath reciprocal multiplication to shared constant divisors and float64, retaining divisions when the reciprocal is non-normal.

  • YDSP: specialize finite positive constant-base powers as scaled exponentials under fastMath, preserving general power calls in strict mode.

  • YDSP: rematerialize constants used between math calls, reducing call-crossing register pressure in kernels such as the compressor.

  • YDSP: share identical immutable entry constants and coefficients across blocks after loop hoisting, reducing duplicate live values on both native targets.

  • YDSP: use encodable integer immediates for fused comparisons on ARM64 and x64, reducing constant-register pressure in delay-bank kernels.

  • YDSP: fold fused-subtraction write-backs in the shared optimizer and preserve overlapping multiply operands in x64 lowering.

  • YDSP: preserve local initializer snapshots when their source state or local is subsequently assigned.

  • YDSP: fold temporary state-copy chains in the shared optimizer, with coverage for saved values and sample history across blocks.

  • YDSP: share operand lowering between ARM64 and x64, including x64 fused product subtraction and width-correct integer immediates; add emission coverage for both architectures.

  • YDSP: contract product-minus-addend expressions and fold them to ARM64 fnmsub; encode small integer offsets and low-bit masks as ARM64 immediates.

  • YDSP: eliminate unconditional loop backedges through comparison-only headers in native codegen; add empty-block coverage and small-kernel block-size benchmarks.

  • YDSP: share compatible input delay taps in one masked history ring, prefer FMA contraction on longer feedback paths, and contract the right-hand product of product differences.

  • YDSP: use indexed optimizer lookups, worklist dead-code elimination and early cleanup convergence; reuse eligible stream scratch by lifetime, simplify delay wrapping, and preserve strict subtraction by negative zero. Add scratch-size reporting, opt-in allocation assertions and compile/dense-event benchmarks.

  • YDSP: retain only process(const YdspProcessRequest&); migrate examples, tests and benchmarks from positional overloads.

  • YDSP: validate processing requests before mutation, add span-based YdspProcessRequest, return preparation validation failures via Result, and explicitly reject sample-accurate automation for rate-converted nodes. Slot APIs use BySlot names to avoid ambiguity with YUP strings.

  • YDSP: parameter/meter getters now read atomic block snapshots, with slot-based access and bounded single-consumer parameter draining. Delayed events survive multiple variable-sized blocks; exact-boundary events belong to the next block. Event dispatch uses ordered cursors while preserving equal-offset precedence.

  • YDSP: emit ARM64 and x64 kernels independently of the compiler host, including WebAssembly; version 2 bundles store native/WebAssembly artifacts and their source closure, resolve helpers at load time, and instantiate without disk imports or machine-code regeneration. Version 1 bundles must be regenerated.

Added

  • YDSP: fmsubF lowers to a single fused fmsub on AArch64 instead of an fmul/fsub pair. FMSUB d, n, m, a computes a - n * m with one rounding, so this removes an instruction, a virtual register and a link from the loop-carried dependency chain of every contracted c - a * b (the ladder filter's in - fb * z4, the wave shaper's 1.0 - env * 0.5). It also settles an inconsistency: lowerFusedMultiplyAdd already gives a target without the instruction one rounding through its float64 expansion, so the FMA-capable target was the less accurate one. Pinned by a numeric test rather than left implicit.

  • Tests: yup_YdspBenchmarkTests gains nine real-life effect benchmarks - a fractionally addressed feedback echo, an LFO chorus, a Freeverb eight-comb/four-allpass reverb, a dB-domain bus compressor, a tanh drive distortion with tone control, a TPT state-variable low-pass, a six-stage phaser, a Karplus-Strong pluck and a 12:1 sample-hold/bitcrush lo-fi processor. Each compares the JIT kernel against a hand-written C++ routine over the standard benchmark length, across the four optimisation policies, and under the same loose parity guard the other shapes use (checksum where the loop stays libm-free, relative magnitude where it does not). The five effect shapes mirror the shipped fx/ example processors with their UI annotations stripped.

  • YDSP: loop-invariant code motion now places a hoisted instruction as early in the sample loop's preheader (the kernel entry block) as its own operands allow - right after the last preheader instruction that defines one of them, never before. An invariant that reads no scalar state (a hoisted tan/exp coefficient) therefore lands ahead of the entry block's scalar-state loads and is no longer crossed by the loop's register-promoted state, which the register allocator previously parked on the stack for the whole block loop; the TPT state-variable filter benchmark dropped from ~1.9x to ~1.25x of its hand-written C++ reference as a result. An invariant that consumes a state register (a voice body's env * gain) still lands after the load that defines it.

  • YDSP: the native backends fuse a comparison into the branch that consumes it. A branchIf whose condition is a single-use integer or float comparison now re-emits that comparison into the condition flags at the terminator and branches on them (jcc on x86-64, b.cond on AArch64), instead of materialising a 0/1 register with setcc/cset and testing it - every sample-loop header and data branch drops the register round trip.

  • YDSP: values live across a call now survive it in callee-saved registers instead of round-tripping through the stack at every call. The bundled AsmJit (thirdparty/asmjit_library) previously spilled unconditionally any value sitting in a register a call clobbers whenever that value's live range crossed a basic-block boundary - which meant a per-sample libm call (sin/tanh/log/pow in an effect chain) stored and reloaded the loop's state, stream pointers and coefficients on every sample. bin_pack now detects call-crossing values by live-span containment over the call sites and gives them preserved homes up front, packing them ahead of any call-free value so none can claim the last preserved register as a fallback first (x19-x28 / d8-d15 on AArch64), and the local allocator only parks a value in a preserved register at a call when one is free, spilling only on overflow. Against hand-written C++ references the chorus benchmark dropped from ~2.4x to ~1.1x, distortion from ~1.6x to ~1.12x, tanh shaper to ~0.97x, and the phaser now runs at ~0.68x. The changes are small local patches over upstream AsmJit, marked at the sites in core/ra_pass.cpp and core/ra_local.cpp.

  • YDSP: scalar libm function addresses are materialized once per kernel instead of at every call site. A per-sample sin/tanh/pow used to rebuild its 64-bit target register (movz + two movks) before each register-indirect branch; the codegen now pre-scans the IR, emits one address load per distinct function at kernel entry (materializeScalarLibmTargets) and lets every call site branch through that shared register, which the allocator keeps in a callee-saved slot.

  • YDSP: foldStateWriteBacks now folds integer state write-backs too, not just float chains. A per-sample wp = wp + 1 ring-pointer update used to be a fresh addI plus a movI write-back into the loop-carried register every sample - exactly the float fmov round-trip the pass already removed - and so did the sample loop's own induction move. Integer arithmetic, bitwise and select producers now write the carried register in place (wp = wp + 1 is one add), dropping one instruction per sample off the state chain in the delay-line, chorus, karplus and lo-fi shapes. The canonical write-back move of a constant-bound loop's induction is preserved (fusion, unrolling and the vectoriser pattern-match the loop body on it before the fold runs); the fold is what removes it after unrolling.

  • YDSP: the vectoriser now accepts blockSize - k (and blockSize + k) loop bounds instead of rejecting them as unsupportedLoopBound. They are runtime bounds like plain blockSize, so the same epilogue machinery widens them to bound & ~(lanes - 1) whole vectors plus a scalar remainder loop, with the stream-access requirement enforced as before; a for i in 0..blockSize - 1 stream loop widens the same way 0..blockSize does.

  • YDSP: the vectoriser now widens loops containing the rounding intrinsics floor, ceil and rint. Each lowers to one native packed instruction (frintm/frintp/frintn on AArch64, roundps on x86), so a block-mode quantiser or bitcrush loop over streams vectorises like any other element-wise loop and stays bit-exact per element. Vector round and copysign remain scalar: round-half-away has no one-instruction packed form on x86 and copysign needs float bitwise sign-mask work both backends do not expose yet.

  • YDSP: if-conversion now speculates an input-stream load when it reads the induction of the stream-length loop that directly encloses it (the sample loop, or a block-mode for i in 0..blockSize loop). A periodic sample-and-hold or downsample branch (if (counter == 0) { held = in; }) becomes straight-line code that loads unconditionally and selects, instead of a branch that mispredicts every Nth sample - the residual gap in the lo-fi 12:1 sample-hold benchmark. State-array and parameter loads stay inside their branch, as do loads under a constant-bound inner loop.

  • YDSP: scalar leaf values that would otherwise be parked around a per-sample libm call are now rematerialized instead. A parameter load or compile-time constant defined before the loop but used only after the loop body's last libm call gets re-defined right after that call (rematerializePostCallLeaves): it reloads once per sample - the same value - and its live range never spans the call, so the register allocator stops emitting the per-sample str/ldr [sp] parking pair and one less value competes for a preserved register. Leaf defs only; expF is not treated as a boundary because fastMath inlines it on AArch64.

  • YDSP: voice banks now work on subgraphs - node v = VoiceChain[8] inlines the whole chain as one runtime voice group, so an effect (filter, delay, reverb, shaper) runs once per voice with its own state instead of once after the summed bank. A banked subgraph must declare an input event and exactly one float32 output stream and contain at least one event-handler member; members must be ordinary single-voice float32 processors (nested banks, per-member rate changes and midi-only members are rejected). Group members run voice-major with per-voice intra-group delay rings, each note event is allocated once per group and fanned to every subscribing member on the same voice, all-sound-off clears the group's slots once and silences every member, and getActiveVoiceCount accepts the bank name. examples/graphics/data/synths/PerVoiceEcho.ydsp is the worked example; the crosstalk test proves two notes through the chain are numerically different from the same effect placed after the mix.

  • YDSP: a state array may now omit its size and let a { ... } initialiser list determine it - state float wavetable[] = { ... } - and the compile-time pseudo function size (...) returns the element count of any array expression: a state array (explicit or inferred; an array of struct instances yields its instance count), a struct array field (size (comb.buf) / size (combs[i].buf)), and, in a block-mode processor, a stream (whose length is the runtime blockSize). A [] state without a non-empty { ... } list, and struct-array states written [], are compile errors with a message naming the fix.

  • YDSP: one state statement may now declare several states of the same type, separated by commas - state float x, y, z;. Each declarator keeps its own array size, initialiser and trailing annotation (state int active [[ role: voiceActivity ]], released; marks only active), and the list is sugar for the equivalent run of single-state statements.

  • Examples: fx.Delay is now a fractionally addressable delay. The read tap is a linear interpolation between the two ring samples either side of the requested delay, and its time parameter is smoothed, so changing the delay time while audio is running glides the tap instead of clicking on whole-sample steps. PerVoiceEcho.ydsp runs one such delay per voice.

  • YDSP: adjacent members of one banked voice chain that are pure per-sample processors now fuse into a single kernel, exactly as plain chains always did. The fused member still runs once per voice under the group's slot table (per-voice state and per-voice parameters are preserved), but the intra-group junction stops being a per-voice scratch round-trip; fusion never crosses a group boundary.

  • YDSP: a banked voice chain now skips a whole voice when every member reports its [[ role: voiceActivity ]] flag asleep - the flag is no longer restricted to processors that declare event handlers, so pure effects (filters, delays, reverbs) can opt in by keeping their flag set until their own tail has died. Skipping freezes the whole voice (per-voice delay rings included); a held voice, a pending event, or block-wide automation/all-sound-off always keep the voice running. Fusion and skipping are complementary: fused effect members cannot carry a flag, so their voices always run.

  • YDSP: the optimizer now performs block-local common-subexpression elimination for pure IR expressions, reducing repeated arithmetic without changing non-SSA state or memory semantics.

  • YDSP: endpoint annotations gain [[ mid: <value> ]] (the value that should sit at the middle of a host slider's travel, exposed as YdspParameterInfo::midValue so a UI can derive a logarithmic skew) and [[ bipolar: true ]] (a range-centered-on-zero flag, exposed as YdspParameterInfo::bipolar, default false). examples/graphics's YDSP Synth Lab applies mid through Slider::setSkewFactorFromMidpoint(); the Analog Saw patch's Cutoff knob demonstrates it.

  • YDSP: introduced the .ydsb bundle API, yup_dsp_compiler host tool, CMake embedding helper, and bundle format documentation.

  • YDSP: fast-math contraction now covers multiply-add and subtract-multiply patterns; the explicit tradeoff is changed rounding. Fixed inline @ delays use a compact increment-and-wrap IR operation for faster native code.

  • YDSP: YdspCompiler now accepts per-compile YdspCompileOptions, providing baseline/automatic/aggressive policies, strict-by-default fastMath, host or portable-target selection and an optional optimisation report. Native bank-loop SIMD now uses the selected target width (SSE2/ASIMD x4 or AVX2 x8), AVX2 emits packed FMA when fast math is enabled, and AVX-width kernels emit vzeroupper on return. AVX-512 is detected but remains disabled until a measured microarchitecture cost model is available. YdspBenchmarkTests now compare automatic code against the scalar baseline on the modal-bank shape, reporting lane width, generated code size and timing while requiring identical strict output.

  • YDSP: the vectoriser now handles the scalar remainder of a trip count, and the per-sample stream loop. A constant-bound loop whose span is not a whole multiple of the lane count peels its leading remainder as straight-line scalar copies and starts the vector loop at a whole number of vectors (a 6-mode bank is two scalar iterations plus one four-lane trip, not six scalar ones); a non-zero constant start is handled the same way over the span, which also fixes an overrun the old divisible-bound-only rule had for such loops. A blockSize-bound loop whose body reads or writes streams at the loop variable (in[i] / out[i], the sample-mode gain/mix shape or a block-mode stream loop) is widened too, with a rolled scalar tail loop after the vector loop, packed stream loads/stores on the native backends and v128.load/v128.store on wasm SIMD, and the reduction fold placed once after the tail. Loops that were already scalar stay scalar: a constant span shorter than one vector, a runtime start, a blockSize-bound state-array bank, or a stream store at a fixed index.

  • YDSP: missed-vectorization diagnostics. YdspVectorizer::run now records one outcome per original loop - widened at the lane count, or the exact reason the loop stayed scalar (shortTripCount, unsupportedWidenedOp, indirectAccess, loopCarriedValue, nonConstantStart, ...) - exposed as YdspKernelReport::loopVectorization (with YdspVectorizationReport::rejectionReasons() for the deduplicated text and YdspVectorizationResult::describe() for one line), and emitted as info diagnostics when YdspCompileOptions::emitOptimizationReport is set, the LLVM -Rpass equivalent for "why is this loop scalar?".

  • YDSP: the WebAssembly backend now lowers the vectorised IR to f32x4 SIMD when the module is compiled with -msimd128 (the emscripten default, which defines __wasm_simd128__). Widened state-bank loops emit v128.load/v128.store, packed f32x4 arithmetic, splats and a shuffle-based horizontal reduction; element-wise work stays bit-exact against the scalar form and the reassociated accumulation keeps the existing tolerance. The automatic tier applies the full native transform set on wasm: vectorisation at four f32x4 lanes, unrolling of the widened loops, and halving of the widened reduction chains (all pure IR passes), while a build without -msimd128 keeps every loop transform off and the scalar-only rejection, now naming the flag. The wasm vector width is fixed at four lanes (128-bit SIMD), so kernels vectorized for AVX2/AVX-512 widths are rejected with a diagnostic rather than miscompiled.

  • MIDI: the WASM (Emscripten) backend of yup_audio_devices is now backed by the Web MIDI API (yup_Midi_wasm.cpp). MidiInput and MidiOutput enumerate, open, start/stop and send to browser MIDI ports; SysEx is requested (sysex: true); hot-plug statechange events update MidiDeviceListConnection listeners; incoming streams are converted through the existing bytestream handlers (so ump::Receiver with MIDI_2_0 works), outgoing ump::View/ump::Packets are converted to MIDI 1.0 bytes, and sends from other threads are proxied to the main thread. createNewDevice() stays unsupported since the Web MIDI API cannot create virtual ports. Note the browser permission is asynchronous: call getAvailableDevices() early and lists populate once the user grants access.

  • Tools: a new modules/yup_dsp_jit/tools/vscode-ydsp extension brings YDSP syntax highlighting, snippets and editing configuration to VSCode, installed with just vscode.

  • Examples: examples/graphics can now be configured to build a single demo instead of the full browser, via -DYUP_EXAMPLE_GRAPHICS_DEMO=<id> (e.g. SpinningCube) in place of the default ALL. A single-demo build compiles out every other demo's code, skips the picker list UI, and only embeds/preloads the rive, lottie, shader and data/synths resources (and the glslang/spirv_cross/spirv_tools shader transpiler) that the selected demo actually needs - useful for small, single-page embeds such as documentation website demos.

  • GUI: yup_audio_gui gains PitchWheelComponent and ModWheelComponent, vertical-drag controls for building a pitch-bend + mod wheel strip next to a MidiKeyboardComponent. The pitch wheel reports a bipolar -1.0..1.0 value and springs back to its default on mouse release by default (settable to hold instead); the mod wheel reports a unipolar 0.0..1.0 value and never springs back. Both are themed as a cylindrical wheel body with a single sliding grip line - not a slider track and thumb. examples/graphics's YDSP Synth Lab now drives pitch bend and CC1 from the two wheels, placed to the left of its keyboard, instead of the two sliders that previously stood in for them in the expression row.

  • YDSP: noteOn handlers can now read e.bendSemitones, the pitch-bend in effect at the moment the note is triggered. Previously a freshly triggered voice only saw the bend as a later pitchBend event, so patches either reset their bend factor to 1.0 on note-on (a key pressed while the wheel was held up started un-bent) or leaked a stale value from a recycled voice. The runtime already computed the value (payload.bend from the MPE note, which in legacy mode is the last wheel position on the channel); it is now exposed on the noteOn shape like pitch/velocity, and mono note-ons carry it too (the mono held-note record now stores the bend). The bundled synth patches replace their bendFactor = 1.0 note-on reset with bendFactor = pow (2.0, e.bendSemitones / 12.0).

  • YDSP: processors can now generate events. output event <name>; declares an emitting channel and emit <shape> (field: expr, ...) -> <name>; sends one, legal in the per-sample process body and in event handlers. A channel's declared name is an identifier, not a shape - it need not equal any of the seven shape names, and one channel may carry several different shapes over its lifetime. In a connection { } block, a node's output event wires to another node's input event (including the polymorphic midi input) or to the graph's own output event boundary, which the host reads back as MIDI; every declared output event must be connected at least once, and an emitted event straddling a block boundary or carrying [[ latency ]] compensation arrives already aligned with its source node's audio. This is what lets a MIDI-only processor - one with no stream endpoints at all, such as the new midi.Arp and midi.Transpose in examples/graphics/data/synths/midi/ - drive an existing, unmodified voice bank by composition alone; ArpPolySine.ydsp and ArpTranspose.ydsp are worked examples, the latter a fully MIDI-only graph with no audio stream anywhere in the patch. A graph is now classified purely by which endpoint kinds it declares (audio-only, MIDI-only, or hybrid), not by a keyword. A MIDI-only node's own note bookkeeping (e.g. midi.Arp's held-note table) is no longer subject to the runtime's ordinary per-voice allocation and stealing: with no stream endpoints, and therefore no per-note "sound" to steal, every event now reaches the node's sole instance directly, so a node like midi.Arp correctly tracks a whole chord rather than just its most recent note. midi.Arp also gained a mode parameter (Up, Down, Up-Down) selecting which direction it steps through the held notes, and restarts its clock on the first note of a phrase so a chord sounds immediately rather than after up to one full 1/rate.

  • YDSP: fixed a bug in the compiler's import-cloning step (cloneStmt()) that dropped an emit statement's shape, target and field list whenever the emitting processor was only ever reached through an import - every existing test that used emit declared its processor inline, so the gap went unnoticed until midi.Arp/midi.Transpose shipped as importable library processors.

  • YDSP: a graph's input event <name>; is no longer a broadcast subscription. Previously any node whose own processor declared input event <name>; with the same name received every event on that port automatically, with no wire in the connection { } block to show it - and which port a node actually heard depended on its declaration order among the graph's inputs, invisible at the call site. A graph input event is now wired exactly like a stream: <name> -> node.event; in a connection { } block, fanning out to as many destinations as are wired and reaching none that aren't; an unconnected graph input event or an unconnected node input event is now a compile error, the same rule already applied to output event. All shipped synth patches (and ArpPolySine.ydsp/ArpTranspose.ydsp) gained the explicit midi -> voices.midi; (or midiIn -> arp.midiIn;) line this requires; PolySine.ydsp moved from the algebra body form to a connection { } block, since the algebra form has no event syntax.

  • YDSP: native transcendentals now lower through the bundled sleef_library (SLEEF, Boost-licensed): widened transcendental loops vectorize to 4-lane calls (u35 under the now-default fastMath, u10 when strict), scalar float32 values stay on libm, fastMath is enabled by default on native targets (wasm stays strict regardless), and an 8-lane AVX2 value splits into two 4-lane calls. Optimizer borrows from the SNEX reference JIT: constant division becomes reciprocal multiplication under fastMath, pow2 modulo of a provably non-negative value becomes a mask, and adjacent same-bound memory-disjoint loops fuse before vectorization; constant math calls fold for the full intrinsic family. Register allocation weights hot values (inductions, widened lanes, stream bases) so they never spill. Benchmarks report a per-policy matrix (baseline strict / + fastMath / host strict / host + fastMath default) with a transcendental-call count per kernel.

  • YDSP native codegen: under the native fastMath default, scalar float32 exp on AArch64 is lowered to a straight-line degree-8 Estrin polynomial (~1 ulp across [-1, 1]) instead of a per-sample libm call plus its register-allocator spill round-trip; arguments beyond |x| = 1 take a rare, predictable fallback branch to libm, and x86-64 keeps the libm call. The exp-envelope benchmark drops from ~4.6 to ~1.8 ns/sample (~1.05x of its C++ reference, down from 2.6x) and the modal bank from ~7.3 to ~5.0 ns/sample (0.74x of C++). The optimizer's block-local CSE now also covers repeated reads of the same input stream slot - input buffers are immutable for the lifetime of a kernel, so the reads are pure - collapsing e.g. the four per-sample in[i] loads of the delay-taps shape to one.

  • YDSP optimizer: scalar state write-backs fold into the value they move. A sample-mode state update lowers to v = op (...) followed by movF y = v (and the builder keeps binding the variable to v); the new foldStateWriteBacks pass rewrites the producer to write the loop-carried register y directly, redirects every in-block use of v to it and drops the move, so the per-sample chain carries no extra register hop. The ladder filter drops from ~10.7 to ~9.7 ns/sample (host + fastMath, now ~1.05x of its C++ reference), the wave folder from ~1.8 to ~1.1 ns/sample (~1.03x of C++), and the exp envelope from ~1.8 to ~1.5 (0.86x).

Changed

  • YDSP: the lexer scans with a UTF-8 character pointer instead of a character index into the source String. String::operator[], length() and substring() each walk the buffer from the head, and the old cursor called them once or twice per source character, which made YdspLexer::tokenize() quadratic in source length; it is now linear. Tokens, line numbers and column numbers are unchanged - columns still count characters, not bytes.

  • Tools: the python stdlib archive generator (used by the tests target) now skips the copy and zip steps when the source bundle and tool configuration are unchanged since the last run, so reconfigures no longer pay for a full stdlib rebuild.

  • Build: yup_add_embedded_binary_resources now regenerates a resource's byte array only when the input file's content actually changed (tracked via an MD5 sidecar next to the generated .inc), so reconfigures no longer re-read and re-serialize large embedded files such as the python stdlib zip.

  • Build: Xcode builds no longer auto-regenerate the project during a build (CMAKE_SUPPRESS_REGENERATION is set for the Xcode generator), so building no longer cancels with "project is being modified while building" when CMake input files change; re-run cmake (e.g. just mac) after editing CMakeLists.txt or .cmake files.

  • YDSP native codegen: kernel prologues now load context pointers only when the generated IR uses the corresponding resource, reducing register pressure and avoidable spills in small sample kernels.

  • Emscripten: the standalone shell now shows a non-blocking hint over the canvas when audio needs a user gesture, reports audio/MIDI availability in the top rail, and keeps activation in the shared AudioWorklet backend for all examples using the shell.

  • GUI: PitchWheelComponent and ModWheelComponent now take a MidiKeyboardState in their constructor, like MidiKeyboardComponent. Both register as state listeners and follow the pitch-wheel / modulation-wheel (CC 1) position of their midi channel (see setMidiChannel()), applied asynchronously on the message thread - updates arriving between message-thread passes are coalesced so only the latest position is applied, and no update is applied while the user is dragging the wheel. The YDSP Synth Lab demo drops its manual atomic bridge (incoming pitch bend / CC1 stored on the MIDI input thread and applied in refreshDisplay()) and lets the wheels follow keyboardState directly.

  • YDSP optimizer: the compiler pass implementations are split out of the monolithic optimiser/passes/yup_YdspPasses.cpp into one file per pass (yup_YdspPassesConstantFolding.cpp, yup_YdspPassesAlgebraicSimplification.cpp, yup_YdspPassesCopyPropagation.cpp, yup_YdspPassesIfConversion.cpp, yup_YdspPassesFullyUnrollBoundedLoops.cpp, yup_YdspPassesSplitWidenedReductionChains.cpp, yup_YdspPassesStoreToLoadForwarding.cpp, yup_YdspPassesDeadCodeElimination.cpp, yup_YdspPassesLoopInvariantCodeMotion.cpp, yup_YdspPassesContractMultiplyAdd.cpp, yup_YdspPassesLowerFusedMultiplyAdd.cpp), with the helpers shared across passes moved to yup_YdspPassesShared.cpp. Pure cut/paste - no behavior or IR change.

  • YDSP: the public API headers are split out of the monolithic compiler/yup_YdspCompiler.h into per-concern files: the runtime graph API (YdspAudioGraph, YdspParameterInfo, YdspExecutionReport, the stream buffers and YdspProcessResult) now lives in runtime/ (yup_YdspAudioGraph.h, yup_YdspTypes.h, yup_YdspExecutionReport.h), and the compiler API (YdspCompiler, YdspDiagnostics, YdspCompileOptions/YdspOptimizationReport, YdspRecursionGuard) in compiler/ (yup_YdspCompiler.h, yup_YdspDiagnostics.h/.cpp, yup_YdspCompileOptions.h, yup_YdspRecursionGuard.h). Pure cut/paste - no behavior or API change.

  • Examples: the YDSP Synth Lab demo (examples/graphics/source/examples/YdspSynths.h) is restyled to match cmake/platforms/emscripten/shell.html's dark theme - the same void/surface/edge/ink/muted/glow palette, and a surface-coloured rail with a hairline edge behind the toolbar/title bars and behind the keyboard row. The toolbar is now one aligned strip: the YUP mark and a bold "YUP!" wordmark, a shortened patch selector, then Performance/Editor tabs, master volume, the oscilloscope and All Notes Off as equal-width slots, so the row reads as a single set of controls rather than mismatched widths; the separate "YDSP Synth Lab" title label was dropped as a redundant second wordmark. The Dump Asm/Dump Wasm button moved out of the toolbar into the editor tab, next to Compile, since it only makes sense there. The row below the toolbar now holds the MIDI input selector alongside the expression sliders, reclaiming the space the old dedicated title/patch/volume row used to waste; and the parameter knob grid expanded into that freed height plus the space the oscilloscope vacated. Parameter cards, meters and the oscilloscope now share the shell's blue accent instead of each having their own. data/logo.png is preloaded unconditionally in the Emscripten build, since main.cpp's own title chrome loads it regardless of which demo is selected and it previously was not.

  • Emscripten: the standalone shell page (cmake/platforms/emscripten/shell.html) has been redesigned around a dark, low-chrome theme keyed to the YUP mark, with the loading indicator doubling as the download progress arc. The canvas display is now a three-way choice - embedded, full window (canvas fills the tab, no scrollbars) and fullscreen - replacing the old resize-canvas and lock-pointer checkboxes, which only fed Module.requestFullscreen and had no observable effect. Fullscreen now reuses the full window layout rather than emscripten's own fullscreen sizing, so the two behave identically. The top rail reports whether the page is cross-origin isolated and how many threads the browser offers, which is what a failing pthreads or audio worklet build needs first. Note that full window scales the existing framebuffer: the app is only redrawn at the new size once SDL_EVENT_WINDOW_RESIZED is forwarded to handleResized in yup_Windowing_sdl.cpp.

  • YDSP: new fma(a,b,c) intrinsic - a * b + c with a single rounding. The compiler has never fused a multiply and an add on its own, deliberately: contraction is a precision liberty, and a patch has to produce identical samples on every backend it can be compiled for. That refusal turned out to be most of what separates the JIT from compiled C++ on per-sample recurrences, where the multiply and the add are consecutive links of the loop-carried chain. A new benchmark variant sizes it by giving the C++ reference a #pragma clang fp contract(off): the ladder filter reads 1.56x against a reference that fuses and 1.05x against one that cannot, so on that shape contraction was essentially the entire gap; the wave folder reads 2.15x and 1.24x. The wave shaper is the control - its recurrence is a multiply and a select, with no add to fuse into - and its two references land together, confirming its ~1.39x is not an FMA story. fma closes the gap without giving up the guarantee: the operation has one defined value, and a target with no fused instruction (wasm, or x86-64 without FMA3) reaches that same value by computing in float64 and rounding once - exact for normal results, since a float32 product is exact in float64 and 2p+2 = 50 bits fit in its 53. So native and wasm still agree; what differs is a patch written with fma against the same patch written with * and +. It is float32 only, because the fallback needs a format one step wider than the operands and none exists above float64 - fma on float64 operands is a compile error rather than a silent per-target difference. The expansion is an IR pass, so neither the wasm backend nor a pre-FMA3 x86 one needs to know the opcode exists; AArch64 lowers it to fmadd, x86-64 with FMA3 to vfmadd213ss. A contractMultiplyAdd pass applies the same rewrite automatically to every a * b + c whose multiply feeds nothing else, so patches get this without being rewritten - and measurably do: with it on, the ladder went 15.2 to 10.3 ns/sample (1.56x to 1.135x) and the wave folder 2.71 to 1.87 (2.15x to 1.645x), both now faster than the reference that cannot fuse (0.77x and 0.94x). The same patches written with fma() by hand measure 0.98x and 0.99x against the automatic form - inside noise, which is the check that the pass finds what a person would and picks the same operand when both are eligible. The wave shaper is the control: its recurrence is a multiply and a select with no add to fuse into, and it did not move (1.39x to 1.41x). The pass runs after the vectoriser and skips anything widened, since there is no portable packed fused form; when both operands of an add are fusable multiplies it picks the one on the recurrence, because fusing the other leaves the loop-carried chain a link longer than it started. Two costs come with it being automatic: on a target with no fused instruction each contracted site is six operations instead of two, which WebAssembly pays throughout and which is not yet measured; and an algorithm that depends on the intermediate product being rounded (a*b - c*d determinants, Kahan summation, Dekker's twoProduct) is changed by it, with no per-expression way to opt out yet.

  • YDSP: a graph is now an arbitrary DAG. Connectivity was "exactly once" on every graph input, graph output, node input and node output stream, which made a YDSP graph a forest of chains: a signal could not be split and rejoined, so dry/wet, parallel multiband, mid/side, a metering tap and summing two sources into one input were all inexpressible. The rule is now at least once on both sides, and beyond that: a source may fan out to any number of destinations, and a destination fed by more than one source sums them, with no mixer node needed. Fan-out costs nothing - generated code never writes through ctx.inputs, so N consumers share one buffer pointer and no new memory is allocated. A new YdspBenchmarkTests shape measures the cost, and the answer is that summing is effectively free. A graph-level dry/wet built from fan-out plus fan-in runs at 6.32 ns/sample against 5.96 for the same patch with dry/wet hand-rolled inside one processor - but most of that 0.43 gap is the second kernel call, not the mix, because the fanned form cannot fuse. A third variant in the same test separates them by fanning out to two separate graph outputs instead of summing into one: same two kernels, same fan-out, but both outputs take the direct-write path so no mix buffer exists. Against that, the mix path costs 0.109 ns/sample (6.32 vs 6.21, a 1.018x ratio) - a memcpy plus one add per sample, which is what it should be. Graph-level dry/wet is therefore the idiomatic form now rather than a luxury, and the remaining ~0.32 ns/sample is simply what a second kernel call costs. Implicit summing requires a float32 or float64 stream, named in the diagnostic when it is not; fan-out has no type restriction, being the same buffer read twice. The four zero-use rejections survive and two of them are load-bearing rather than stylistic: a graph output with no source would have nothing written to it (and the runtime never zeroes an output buffer, so the host would hear its own uninitialised memory, not silence), and a node input with no connection leaves runtimeInputs[s] null for the kernel to dereference. Mixing happens on the input side and the destination owns the mix buffer, which leaves a node's exclusive ownership of its runtimeOutputs - and therefore the polyphonic pre-voice-loop zeroing, clearVoiceSpan and the voice accumulate - untouched; connection 0 memcpys into the buffer and 1..N-1 accumulate, so "written exactly once" is structural rather than an ordering rule to get right. Summation order is the order the edges appear after analysis: deterministic per patch, and nothing more (subgraph inlining and kernel fusion both rebuild the edge list), so it is preserved with a counting sort rather than a std::sort, and documented as something not to depend on. Existing patches are bit-identical: a graph output keeps writing straight into the host buffer when it has exactly one source, that source is a node, the edge carries no delay, and that node output has exactly one destination - which every previously-legal shape satisfies. MasterBus.ydsp is the first patch in the tree to use the feature, gaining a real parallel dry path around its reverb so reverbMix is a graph-level balance instead of a value forwarded into Reverb's own internal mix; three new demo patches showcase it: ParallelRack.ydsp (one voice fanned out to three character paths, summed back), HaasWidener.ydsp (one mono chain fanned out to both graph outputs with an inline delay on the right), and ParallelDrive.ydsp (a clean path summed against an oversampled hard clipper feeding a [[ latency: 32 ]] lookahead limiter, so the 48-sample skew is compensated rather than comb-filtering the blend). The four bundled effects are deliberately left alone as regression baselines.

  • YDSP: automatic plugin delay compensation, and a latency figure to report to the host. Before paths could reconverge, latency misalignment could not be heard, so YDSP had no latency concept at all - zero occurrences of the word in the module. The moment fan-in exists, an oversampled branch summed against a dry one is comb-filtered rather than merely late. The new computeLatencyAndCompensate pass equalises it, and YdspAudioGraph::getLatencySamples() reports what is left (feed it to AudioProcessorBase::setLatencySamples(), which already drives every VST3/CLAP/AU/AUv3/AAX/LV2 wrapper). What makes it correct rather than merely present is which latency it compensates: an artifact, where the sample count is a leaked consequence of an implementation choice the author did not make (an oversampler's group delay; a new processor P [[ latency: N ]] declaration, which only the processor's author can know), as opposed to intentional, where the count is the semantics because the author typed the number (-> [400] ->, x @ 400). edge.delaySamples appears nowhere in the model. The decisive case is a dry/wet delay effect: the wet path's @ is intentional, so both branches have artifact latency 0, nothing is inserted, the dry stays dry, and the patch reports 0 - where a naive "total group delay" model would delay the dry branch, destroying the effect, and tell the host a 500 ms echo was 500 ms of plugin latency. Compensation always lands on the reconvergence edge, never hoisted upstream, so a branch with one incoming edge always gets 0 and a plain chain is untouched. Separate graph outputs are equalised against each other as well, which is forced rather than chosen: every format YUP targets reports one scalar, so a per-output latency vector is unrepresentable and a patch whose L is 16 samples later than its R would be permanently skewed in every host with no diagnostic - while a wanted skew, written -> [16] ->, is never touched. The oversampler's contribution is derived from the same ydspOversamplerSincRadius constant the runtime instantiates its yup::Oversampler with (16 input-rate samples, linear phase, exactly integral for every factor - which is why an integer delay compensates it perfectly) rather than restated in a comment. [[ latency: N ]] is declared in the processor's own sample domain, so a * 4 instance divides by 4 and a non-dividing factor is a hard error naming both numbers; rounding would ship a sub-sample residual inside the one feature whose job is phase alignment. YdspAnalyzedEdge::compensationSamples is kept apart from delaySamples so the latter keeps meaning "what the author wrote", the fusion predicate keeps its meaning and the pass stays idempotent - the compiler sums them at one line. YdspAnalyzedNode::latencySamples is the sole source of truth afterwards, because fuseNodeChains synthesises a processor declaration that would report 0; a fused node's latency is set to the sum of its members'. The pass runs after fusion, so it cannot cost a fusion opportunity. The reported value is a compile-time constant: a YDSP graph is a fixed DAG, so nothing at runtime can reroute it or change an oversampling factor, and a plugin offering an oversampling selector recompiles on the control thread instead - which is also all the formats support, since changing latency needs a restart request rather than a realtime notification.

  • YDSP: the <: (split) and :> (merge) algebra operators are real. Both were tokenised, given their own YdspOperator, parsed at one precedence level with : and then quietly handed to composeSequential - so they behaved as a plain : and the documented arity rules were never applied (and were stated backwards: <: widens, :> narrows). AlgebraValue's port lists became bundles - one set of terminals per channel rather than one terminal - because "channel 0 goes to two places" and "channel 0 is the sum of two places" cannot be said otherwise, and every wire-emitting site became a cross product over them. All three operators now share one composeFanned: a <: b requires a.outArity to divide b.inArity and assigns b's input j to a's output j % a.outArity; a :> b requires b.inArity to divide a.outArity, sends a's output i to b's input i % b.inArity, and sums the collisions; : is the case where the two are equal. The result arity is (a.inArity, b.outArity) throughout. Neither operator needs a relay node, and neither does _: it carries no ports, so _ <: (a , b) and (a , b) :> _ emit zero wires and are handled by regrouping the other operand's bundles, with the fan materialising when the value is later sequenced against a real leaf. An unconstrained _ defaults to arity 1 on the side the operator governs (left of <:, right of :> - the maximum fan) and to the known side's arity otherwise. process = dry <: (Distort , Chorus) :> wet; is now a parallel dry/wet, and compiles to a graph bit-identical to writing its four edges out by hand. This only became implementable once fan-out and summing fan-in existed: a split is a fanned-out edge and a merge is a summed one.

  • YDSP: a subgraph boundary port can now carry a fan on either side. inlineSubgraphs held one int per boundary port, written unconditionally, so a second edge onto sub.in (or out of sub.out) silently replaced the first and one source or destination vanished - with fan-in now analyzable, that would have been silent wrong audio rather than a compile error. Each side is a list, and the splice is the cross product of the sources reaching a producer's boundary and the destinations leaving a consumer's boundary; the nested loop is required for a pass-through edge inside a subgraph, which touches both boundaries at once and which the two previous independent ifs only handled for the 1x1 case. Each resulting edge accumulates the internal edge's own delay plus the parent edge's delay at each end, added per resulting edge rather than hoisted out of the loop. paramAlias/meterAlias stay single ints: a node parameter still takes one writer.

  • YDSP: undersampling (node x = P / 4) is a real decimator. It previously only shortened ctx.numSamples, so the kernel filled blockSize / N samples and the rest of the block kept whatever the previous one left there, while an odd block size dropped the remainder outright - it was briefly rejected outright during this work rather than shipped broken. It now band-limits and decimates the node's inputs, runs the kernel at 1/N, and interpolates its outputs back, reusing the same yup::Oversampler windowed-sinc machinery *N does (factors 2, 4 and 8; float32 streams only, like *N). Two Oversampler instances are needed rather than one, driven from opposite ends: upsample() resizes the shared oversampled buffer, so a single instance's direct writes for decimation would fight with it. That in turn needed downsample() to stop requiring a preceding upsample() - it now derives the oversampled length from its own numSamples * OversampleFactor (which the assert it replaced already proved equal) and its capacity checks are against what prepare() allocated, so the class is usable in either direction with no new API. getOversampledChannelData() is bounded by the allocation for the same reason: a decimate-first caller has to fill that buffer, and would otherwise have nowhere to write. "Is a block pending" remains getOversampledNumSamples()'s job, which is the accessor designed for it. Four new OversamplerTest cases cover the direction (accessor availability after prepare(), standalone anti-aliased decimation, a decimate-then-interpolate round trip in the two-instance shape, and continuity across block boundaries). The block-size problem *N does not have - /N can only consume whole groups of N samples, so any block size that is not a multiple of N leaves a remainder - is solved with two carry FIFOs whose counts are invariant at N - 1 between them, which is exactly the surplus needed to fill every block; priming the output side with N - 1 zeros is what establishes it, and costs N - 1 samples of latency. Total /N latency is 16 * N + (N - 1) graph-rate samples (the resampler's 2 * SincRadius is in the node's domain, hence the multiplication), all of it artifact latency that the new delay compensation removes automatically - which is exactly what the artifact/intentional rule predicted for a decimator before one existed. Relatedly, sampleRate inside a rate-changed kernel now reports the rate that kernel is actually running at rather than the graph's; it was previously always the graph's, so an oversampled processor deriving any coefficient from it (which is the only reason to ask) ran those coefficients N times too slow, and an undersampled one would have run them N times too fast. The new ControlRateWah.ydsp demo is built on both: an envelope follower - control-rate work by nature - runs at / 8 while the audio path stays full-rate, and the 135-sample skew where the two reconverge is compensated away, which is what stops the filter opening ~2.8 ms after the transient that opened it.

  • YDSP: the feedback-cycle diagnostic no longer implies an inline delay would fix it. rebuildTopoOrder rejects every cycle but said "feedback cycle without a delay", which read as a promise that -> [1] -> would break one. It now says plainly that cycles are unsupported in this version and that a delay does not break them.

  • YDSP (breaking): four syntax changes. Function parameters are now name: type - func noteToFreq (pitch: float) : float { ... } - and the old (float pitch) form no longer parses. Imports use dotted module paths: import fx.Delay maps to the file fx/Delay.ydsp and forces access as Delay.* (import X.Y.Z as W forces W.*); importing two different files that would share a namespace is a compile error suggesting an as alias, and a circular import is now a hard error (it was a warning). A file that declares only top-level funcs is a library, imported like anything else and called as ns.funcName (...). Events are now named channels: input event <name>; at processor and graph scope, any name and any number (like streams), each carrying all seven shapes; a handler selects the shape with event midi (e: noteOn) { ... } (the old per-shape input event noteOn; / event noteOn (e) pair is gone). Each graph event input is a separate event stream - a node subscribes to the graph input matching its processor's input event name, and process() gained a span-based overload (yup::Span<const yup::MidiBuffer*>) plus getEventInputCount()/getEventInputName() to feed them individually; the single-buffer overload feeds the first input.

  • YDSP: YdspCompiler::compile() takes an optional ThreadPool* (third parameter). When one is passed, imported files are read, lexed and parsed in parallel on that pool; the merge stays single-threaded, the results are identical to the sequential path, and the pool remains caller-owned - the compiler never removes jobs it did not add.

  • YDSP: a graph parameter can now drive more than one node parameter. It was capped at one, which made a wrapper graph unable to do the one thing it exists for - MasterBus could not forward a single mix to both its compressor and its reverb, and the diagnostic read as an arbitrary limit. The cap is gone from validateConnectivity (a node parameter still takes at most one writer, which is what keeps the aliasing unambiguous), and Pimpl::paramSlotToNode became a list per slot so sample-accurate automation is queued on every node the slot drives instead of only the last one wired - previously a second edge silently overwrote the first's routing, so automating that parameter moved one stage and left the other at its old value. The per-block copy path needed nothing: node.paramCopies was already per-node.

  • YDSP: graphs can now be composed. A program may declare several graph blocks - the entry point is the one annotated graph Name [[ main ]] { ... }, or, when the program declares just one graph of its own, that one - and a node may instantiate a graph instead of a processor, so a reusable chunk of signal flow (an effects chain, a voice architecture) gets a name and a parameter surface of its own. import now merges the imported file's graphs alongside its processors under the same namespace prefix, so node master = fx.MasterBus; works across files; an imported graph is always a subgraph, whatever [[ main ]] it carries for its own file. A subgraph is compiled away: the semantic analyzer splices its nodes and edges into the graph that uses it, so the optimiser, all three backends and the whole runtime keep seeing one flat graph and nothing past analysis changed. Inner nodes take the instance path as a name prefix (master.compressor.threshold), a subgraph's input value forwards to whatever node parameter it drives (the node's override list sets it; a parameter edge onto the subgraph node aliases straight through), an output value forwards outwards the same way, and inline delays on both sides of the boundary add up into one edge. Graphs are inlined innermost-first in dependency order, so nesting is unbounded and a graph that reaches itself is a compile error naming the loop. The three things that are runtime mechanisms rather than wiring are rejected on a subgraph node with an explicit reason: a voice bank (Sub[16]) needs one processor with one float32 output stream for the per-voice summing path, over/undersampling (Sub * 4) is a per-processor rate change, and graph-scope input event midi belongs to the entry point. A subgraph parameter that drives no node inside it warns rather than silently doing nothing. The demo's "Analog Saw" patch uses the new examples/graphics/data/synths/fx/MasterBus.ydsp, a graph composing the existing Compressor and Reverb processors.

  • YDSP: a voice bank can now skip idle voices. A processor nominates one state int scalar as its activity flag with the new state annotation state int active [[ role: voiceActivity ]];, and the runtime skips a voice whose flag reads 0 and whose key is not held - no kernel call, no output accumulation - so a Voice[16] bank playing a three-note chord pays for three voices instead of sixteen, and an idle one pays for none. The patch declares completion itself (it knows its own delay lines, high-Q filters and release curves) by testing the amplitude it actually emits; the runtime supplies the held check, so a patch that clears its flag too early wastes CPU on a held voice rather than going silently mute. The flag is read once per voice per block and re-checked after an event handler runs, so a mid-block note-on wakes the voice for the rest of the block, and a voice is never put back to sleep mid-block. Entirely opt-in: a processor that declares no flag keeps byte-identical behaviour. An unknown role: value, a non-int or array-typed flag, a second flag, an initialiser, or the annotation on a processor with no event handlers are all compile errors. The shipped PolySine, AnalogSaw, FMBell, WobbleLead, PulseBass and ElectricPiano patches opt in. New YdspAudioGraph::getActiveVoiceCount ("nodeName") reports how many voices will run on the next block, sharing the scheduler's predicate exactly.

  • YDSP: new kernel-fusion pass. A chain of nodes that only feed each other - x : Osc : Filter : Gain : y, or the same wiring as a connection block - is compiled into one kernel instead of three, so the two blockSize scratch buffers the intermediates round-tripped through become registers. Measured on a new YdspBenchmarkTests shape that compares a three-stage chain against a single processor computing the same thing: the chain went from 5.68 to 2.68 ns/sample, a 2.1x speedup, landing on the hand-fused processor's 2.70 ns. The shape had measured a 2.05x gap before the pass existed - the win had been asserted from first principles and never measured, which is why the item stayed open - so the ratio now sitting at ~1.0 is what the shape guards. The pass works off the analyzed graph, so the : algebra and an explicit connection block fuse alike, and it runs after subgraph inlining, so a chain assembled out of a subgraph's nodes fuses exactly like a hand-written one. A link is fused only when the producer feeds nothing else and the consumer is fed by nothing else - an intermediate tapped to a second destination is observable and stays a real buffer - the connection carries no inline delay ([N], which only the runtime's delay buffer can provide), and both ends are per-sample single-in/single-out kernels that are not voice banks, rate-changed or event-driven. The fused processor is synthesized as an ordinary processor and analyzed through the normal path, so the IR builder, all three codegens and the runtime scheduler see nothing new; every name a member declares is rewritten with a per-member prefix, since a block statement does not open a scope in this language and two members' locals would otherwise collide. Host-visible parameter and meter names are preserved - first.gain and first.level still resolve after the node they belonged to has been fused away - via a new per-node public-name override, a graph parameter aliased onto a member's parameter still drives it, and a member's meter wired to a graph output value is rerouted onto the fused node so it keeps reporting. Once absorbed, a member's own kernel is dropped rather than JIT-compiled as unreachable machine code that getExecutionReport() would list as something the patch runs; only what fusion orphaned is dropped, so a processor still instantiated on another branch survives. YdspDiagnostics gained mark()/rollbackTo() so that a synthesized processor which fails to analyze is silently declined rather than failing the user's compile against source they never wrote.

  • YDSP optimizer: new bounded-loop vectoriser. On the native backends a constant-bound loop over parallel state float arrays - an oscillator, partial or modal bank - is widened to four float32 lanes, so a 16-mode loop runs four packed iterations instead of sixteen scalar ones. It is an IR → IR pass rather than codegen widening, so both native targets share it and it is testable without a JIT; a widened value simply carries lanes > 1 and the existing arithmetic opcodes are reused (addF at four lanes is addps / fadd v.4s), with only vsplat and vreduceAddF added for the two places where the operand and result lane counts differ. Only trip counts that are a whole multiple of four are widened, so there is no scalar epilogue and no block is created, emptied or reordered - the CFG-linear layout both backends recover regions from is preserved by construction. The pass runs last, after loop-invariant code motion has hoisted the loop's shared work (a bank's exp-derived drive term) into the preheader, where it stays scalar and is broadcast once. An accumulation (sum = sum + z[i]) becomes a vector accumulator folded once on the way out, which is what breaks the serial dependency between iterations - and it reassociates, so a widened reduction is not bit-exact while element-wise widening is. A widened body also sinks its loop-invariant scalar prelude into the preheader: in read inside a loop lands there as a loadInput and stays, because loop-invariant code motion never hoists a load, but in a widened body it provably cannot be invalidated - storeOutput disqualifies the loop outright, so nothing in it can write stream memory - and the index is already known to be loop-invariant. The load, everything pure that depends only on it, and the broadcast they feed therefore run once per loop entry instead of once per element. State-array loads are excluded: z[j] at a loop-invariant j is the same element as the widened z[i] store on whichever iteration i == j. Measured against the hand-written C++ equivalents in YdspBenchmarkTests on AArch64: the 32-partial harmonic bank went from 3.20x slower to 0.82x - faster than the compiled reference - and the 16-mode modal bank from 7.51x to 2.35x, of which the last 0.65x came from that hoist. Shapes with no qualifying loop - a delay line, a ladder filter, a wave shaper - are unchanged. A loop is left scalar unless every array element is reached through the loop variable itself, the body is a single block, the loop variable is used for nothing but indexing, and nothing in it touches a stream, a scalar state slot, a ' / @ / smooth slot, a transcendental, a comparison or a select on a widened value. YdspKernelReport gained vectorized and vectorWidth. The wasm backend has no 0xFD-prefix opcode family, so that path stays scalar and rejects a widened kernel outright rather than emitting scalar code for packed values.

  • YDSP optimizer: new ifConversion pass. A short, else-less if whose body is a single side-effect-free block becomes straight-line code plus one select per assignment, turning a data-dependent branch in a sample loop into a conditional move - a wavefolder taking two such branches per sample was paying an unpredictable branch for each while compiled code emitted fcsel. The body is only speculated when every instruction can be executed unconditionally: no memory access (a guarded index may be out of range when the guard is false), no call, and at most eight instructions, so both paths together stay cheaper than a mispredict. The assignments keep their original order, so an instruction that reads a register the body already assigned still sees the conditionally-updated value. No block is added or removed - the condition block absorbs the body and both fall through - so every block index, loop bound and wasm structured region is unchanged.

  • YDSP optimizer: loopInvariantCodeMotion now hoists into each loop's own preheader instead of only the function entry block, using real dominator sets rather than an "is it in block 0" test. Work that is invariant across an inner loop but varies per sample - a filter coefficient derived from an envelope, say - previously could not be hoisted anywhere: the entry block would freeze it at its pre-first-sample value, and there was nowhere else to put it. A benchmark bank of 16 modes sharing one exp-derived drive term was evaluating that exp once per mode instead of once per sample. No block is inserted: construction is CFG-linear, so every loop is already preceded by its preheader (the function prologue for the sample loop, the induction initialiser's block for a for), which leaves the block layout the wasm backend recovers regions from untouched. The single-definition guard on the hoisted instruction's own result is kept - that is what pins a promoted state register, deliberately written by both the prologue load and the per-sample move, inside the sample loop.

  • YDSP optimizer: new storeToLoadForwarding pass. A state-array read that follows a write to the same element becomes a move from the value just stored, so a[i] = x; ... = a[i]; no longer round-trips through memory - in a loop over parallel arrays (the ElectricPiano partial bank reads oscI[i] back immediately after writing it) that also removes a store-to-load-forwarding stall from the accumulation chain. The rewrite replaces one definition with an identical value, so it needs no SSA property. It forwards only within a block, and only when the region, the index value and the stored value are all provably unchanged in between; an intervening array store blocks it unless that store writes a different region through the same index and element width, where the two addresses differ by their region bases alone and so cannot alias for any index. The @ delay ring is unaffected: it writes and reads through different indices.

  • YDSP codegen: a comparison whose only consumer is a select now feeds it through the condition flags instead of a 0/1 general-purpose register. select (a > b, x, y) was emitting fcmp, cset, cmp, fcsel on AArch64 (and comiss, setcc, movzx, test, cmov on x86-64) where two instructions suffice - and the cset/cmp round-trip added two links to a dependency chain that, in a per-sample recurrence such as an envelope follower, is the critical path. The comparison is re-emitted next to the select rather than moved, so nothing can clobber the flags in between; the fusion is skipped unless the comparison is used only by that select, both are in the same block, and neither of the comparison's operands is rewritten between them. Candidates are identified by instruction position rather than by result value id, so a select written by if-conversion - whose result is a mutable local, and therefore defined more than once - fuses like any other.

  • YDSP codegen (AArch64): state-array accesses no longer rebuild their region's base address on every access. AArch64 has no base + index + offset addressing mode, so reaching array[i] means materialising stateArrays + regionBase in a register first - which was emitted inline at each access as a mov plus an add (two extra instructions and two extra virtual registers), even though the region base is a compile-time constant. One register per distinct region is now computed once in the prologue. In a loop over parallel state arrays this was the bulk of the emitted code: the benchmark's 32-partial harmonic bank performs eight array accesses per partial, so it was emitting roughly twenty-four instructions per partial where eight are needed. x86-64 folds the same address into one addressing mode and is unaffected.

  • YDSP codegen: the generated sample loop no longer pays a function call, a redundant memory round-trip or an unpredictable branch per sample. The @ delay wrap lowers to a new wrapI opcode (compare + cmov/csel/select) instead of modI, which on x86-64 was a call into a helper for every delay tap on every sample; scalar state and the hidden ' / @ / smooth slots are loaded once in the kernel prologue and written back once in the epilogue instead of round-tripping through memory every sample; stream channel pointers are hoisted out of the loop rather than re-derived at every access; float constants are read from AsmJit's constant pool (and negF / absF from a sign mask there) instead of being materialised through a general-purpose register, which also frees the registers LICM used to pin across the whole kernel; select is branchless on both targets (fcsel/csel on AArch64, cmov plus an and/andn/or blend on x86-64); AArch64 comparisons use fcmp/cmp + cset instead of a five-instruction branch diamond; and a terminator whose target is the next block falls through instead of emitting a jump. -0.0 is the only observable behaviour change: negF now flips the sign bit, so -(+0.0) is -0.0 on x86-64 as it already was on AArch64.

  • YDSP tests: new YdspBenchmarkTests time six discriminating patch shapes (@ delay taps, a four-pole ladder, a 32-partial harmonic bank, compare + select, a modal bank with work invariant across its inner loop, and a wavefolder with data-dependent if/else) against hand-written C++ equivalents and print a ns/sample ratio for each, with a loose regression guard rather than a threshold assert. Two further shapes compare the JIT against itself to size a specific opportunity rather than to defend a number: the modal bank with and without its accumulation, whose delta is what a widened reduction costs (four serially dependent vector adds plus the horizontal fold), and three chained sample-mode nodes against one processor computing the same thing, whose gap is the headroom a kernel-fusion pass could recover - the win fusion was assumed to have, never measured. The reduction shape has since produced a more useful result than the one it was built for: removing the accumulation makes the kernel reproducibly slower (14.4 vs 20.8 ns/sample, eight fewer instructions, both loops confirmed widened), and a kernel whose runtime moves inversely with its instruction count is not issue-bound - which rules out loop overhead, and with it a full-unroll pass, as the modal bank's remaining gap. It is also the only shape in the suite with a ~30% run-to-run spread rather than under 3%, and the only one calling a transcendental inside the sample loop, so the shape now prints stack-traffic and call counts from the generated listing alongside the instruction count. Separately, the wave folder shape now also times a reference compiled with floating-point contraction disabled: last = last * 0.5f + y * 0.5f is one fmadd on AArch64 by default, which is one rounding where the source asks for two - a precision liberty YDSP does not take, because a patch must produce identical samples on every backend - so the ordinary c++ row is not running the arithmetic the JIT runs, and on a shape whose critical path is exactly one multiply-add per sample that is most of the ratio: 2.26x against the ordinary reference, 1.29x against one that also cannot fuse. The two references produce bit-identical output, because * 0.5f is exact and fusing its rounding changes nothing numerically - which is the check that the pragma took effect rather than being ignored.

  • YDSP optimizer: new full-unroll pass for constant-trip-count loops. A for loop whose trip count is known at compile time is written out in full into its preheader, which deletes the bound compare and the back edge from every iteration. It runs after the vectoriser and so unrolls at the widened trip count - a 16-mode bank at four lanes becomes four copies of a four-lane body, not sixteen of a scalar one - and running it before would have left nothing loop-shaped to widen. No block is added, removed or reordered: the header and body are emptied and left falling through, the same trick ifConversion uses, so every block index, every loop bound and the region layout the wasm backend recovers are untouched; the loop's entry in fn.loops survives so the report still answers "how many iterations could this run" with the worst case rather than 1, marked unrolled so no backend wraps a region around blocks that no longer branch. The body is copied verbatim, increment and all, rather than substituting a constant index per copy: that makes the unrolled code the identical instruction sequence the loop executed, so it is bit-exact by construction and needs no reasoning about a non-SSA IR. It declines a one-shot init kernel, a nested or runtime-bound loop, and any loop whose copies would exceed 256 instructions or 32 trips (256 admits a 32-partial bank at ~190 instructions, but that shape then moved by ~2%, inside its own spread, so the step up from 128 is unproven), and like the vectoriser it is off for wasm - a module is downloaded and parsed before it runs, and the browser's own engine re-optimises the loop regardless, so trading module size for branches is the wrong way round there. This also surfaced a latent trap in the AArch64 backend: vectorIndexRegs memoizes index << 2 keyed by value id and cleared per block, which is only sound while a value id is written at most once per block - an unrolled loop writes its induction variable once per copy, and the stale scaled register would have addressed the wrong element with no diagnostic. A new onValueRedefined hook invalidates the memo at every write. Measured on the modal bank, whose 16-mode loop widens to four lanes and unrolls to four copies: 2.17x slower than the C++ reference down to 1.21x (14.3 to 8.3 ns/sample). The larger result is what it did to the variance - that shape had been the only one in the suite swinging ~30% run to run, wide enough that it had reported a kernel with eight fewer instructions as the slower one and sent two investigations (spill traffic around the per-sample exp call, then the kernel not being issue-bound) after explanations that were not there. Unrolled, the spread is under 3%, and the shape finally produces the number it was written for: the widened reduction costs ~1.9 ns/sample, about a fifth of the kernel. YdspKernelReport gained unrolled, since boundedIterationCount deliberately reports the same worst case either way and is therefore no help in telling the two forms apart.

  • YDSP optimizer: an unrolled bank's widened accumulator is now turned into a reduction tree. Unrolling leaves acc = acc + x once per copy and each add waits on the one before it; sending the odd copies to a second accumulator and adding the two at the end halves that serial depth for the cost of one instruction (the first odd link becomes a move rather than an add, since the second accumulator has no zero to start from, so the add count is unchanged and the combine puts one back). Halving can repeat: the scan resumes just inside the chain it rewrote and picks up the suffix still on the accumulator, so an eight-link chain halves twice and ends at depth four. That is short of what restarting the scan would reach, which is untried. It applies only to a widened accumulator: the vectoriser has already re-associated that sum - lane j sums elements j, j+4, j+8 … and the lanes are folded pairwise - and that is documented as not bit-exact, so this stays inside a licence the language already takes, while a scalar accumulator carries none and is left exactly as written. Measured on ModalBankReductionCost, which times the same bank with and without its accumulation: the reduction went from 1.93 ns/sample unsplit to 0.98 and 1.53 across two split runs, and the modal bank overall from ~1.24x to 1.11-1.20x of its C++ reference. Ranges rather than figures: the direction is unambiguous and the size is not, which is also why the switch exists. That was not the expected result - the reduction looked closer to throughput-bound than latency-bound (six extra instructions, ~6.4 cycles), which is why setReductionSplittingEnabled is a switch separate from the unroller it depends on. The repeated halving was not designed either: it was found by reading an instruction count that had grown by two combines rather than one, and kept because the largest bank in the benchmark posted its best figure with it. YdspKernelReport gained reductionSplit, so a patch can be asked whether its reduction was shortened without inferring it from a timing.

Fixed

  • GUI: Component::hitTest() is consulted again when routing mouse events. Component::findComponentAtForMouseEvent(), which replaced the windowing layer's own traversal, tested only visibility and bounds, so a component that carved out part of its rectangle - a round knob keeping its corners transparent, a control with a padded margin - received moves, enters, exits, clicks and wheel events over the area it rejects, and whatever sits behind it received none.

  • Thirdparty (sleef_library): the SIMD translation units reference Sleef_x86CpuID from their exported dispatch queries (Sleef_getIntd2/Sleef_getIntf4, the DORENAME-renamed xgetInt/xgetIntf), but the definition in SLEEF's src/common/common.c was never compiled into the module, so an x86-64 build failed to link (LNK2019 for Sleef_x86CpuID in sleef_library_simddp.obj/sleef_library_simdsp.obj on the Windows CI; ARM targets were unaffected because their helper header has no CPUID probe). The module now compiles upstream/src/common/common.c as its own translation unit (sleef_library_common.c), as SLEEF's own build does.

  • YDSP: the four-lane x86 float comparison was miscompiled. vectorFloatCompare emitted legacy cmpps with three operands and five-bit AVX predicates, but CMPPS is two-operand (dst = cmp (dst, src)) and decodes only three bits - so the second source was dropped and the destination was read without being seeded. The path now seeds its destination and uses the legacy predicates, emitting b < a for a > b since the three-bit set has no greater-than form. Reachable on a pre-AVX2 host or under targetPolicy = baseline with baselineTarget = sse2; no asmjit validation is enabled in this module, so nothing rejected it at finalize().

  • YDSP: INT_MIN / -1 now yields 0 on AArch64, as it already did on x86-64 and wasm. The pair has no representable quotient; the x86 helpers had always guarded it, the wasm lowering guards it explicitly (div_s would otherwise trap), and the AArch64 lowering - whose comment claimed parity with x86 - guarded only the zero divisor and let SDIV wrap the quotient to INT_MIN. The same patch therefore produced a different value on Apple silicon than on an x86 host. CMN b, #1 gates the INT_MIN compare so the common path pays one extra compare, on an operation the @ delay wrap no longer uses at all.

  • YDSP: integer division and modulo on x86-64 no longer make an out-of-line helper call per occurrence. emitIntDivision now inlines cdq/cqo plus idiv behind the same zero-divisor guard the AArch64 lowering uses, so both targets keep returning 0 for a zero divisor, plus a guard for INT_MIN / -1 - the one pair with no representable quotient, which raises #DE on x86 rather than wrapping. The orphaned divInvoke member and the four yupDspIdiv/yupDspImod helpers go with it.

  • YDSP optimizer: splitWidenedReductionChains now snapshots each link's addend to a fresh value at the link itself before building the balanced tree. Loop unrolling duplicates a body verbatim, so one value id is redefined once per copy; the tree is emitted after the whole chain, so every leaf previously resolved to the last copy's definition and was summed repeatedly - the ElectricPiano harmonic bank summed only its final 4-lane product, making the patch sound as a thin ~0.1 s high-frequency click under the automatic tier (scalar and baseline tiers were correct, which is why only the widened demo patch was affected). The new YdspElectricPianoTests::SustainsPastTheAttackClick and the YdspExamplePatchTests ElectricPiano sustain tests are the regression guard.

  • YDSP analysis: each top-level body (process, init) is now its own local scope, so a same-named local in both no longer reports a spurious duplicate-symbol error; a node written P * 0 / P / 0 is rejected with a diagnostic instead of reaching the latency math with a zero rate factor.

  • YDSP optimizer: if (c) { y = a; } else { y = b; } where each arm moves into the same value is now fused into one select (the wavefolder / clipper shape), removing the branch. pow (x, 2.0) becomes x * x under fast math (strict keeps the intrinsic, since libm pow is not bit-equal to a multiply). Graph meter reads use compile-time byte offsets instead of re-scanning the meter type list per getter. A spelled fma()/fmsub() no longer disqualifies an otherwise widenable loop on targets with a packed fused multiply-add (x86 AVX2+FMA and AArch64), where the widened op lowers to vfmadd213ps/fmla.

  • YDSP runtime: output events carried across a block boundary are stored relative to the next block's start and delivered directly, so an event that straddles into a differently-sized block (64 then 32 samples) still lands at its exact sample instead of being unwound against the wrong block size. The host-side automation and MIDI sample offsets are also clamped to [0, blockSize - 1] rather than the inclusive blockSize.

  • YDSP runtime: a polyphonic node's shared per-voice scratch is cleared for each awake voice before it renders (previously once per block), so a voice whose kernel reads its own output (out = out + 1.0) or writes it conditionally no longer sees the previous voice's samples - two held feedback notes now sum to exactly 2.0 per sample instead of leaking 3.0.

  • YDSP runtime: process() no longer silently truncates a request larger than the prepared block size; it returns the new YdspProcessResult::blockTooLarge and leaves the graph untouched. reset() and prepare() now clear the dropped-event counters (the host-visible droppedEventCount and every node's output-event queue count), which previously grew monotonically for the life of the graph.

  • YDSP runtime: accumulateStream refuses (rather than reinterpreting as float32) when a non-float stream ever reaches the summing fan-in path; the analyzer already rejects such graphs.

  • YDSP wasm backend: integer div/mod now also guards INT_MIN / -1 (wasm div_s/rem_s traps on it, like a zero divisor), mirroring the x86/AArch64 helpers. The emscripten kernel registrar checks its registry before compiling, so prewarmKernels() no longer recompiles an already-registered module.

  • YDSP runtime: setParameter/setDoubleParameter/setIntParameter no longer memcpy straight into the params block the audio thread reads; they push (slot, raw bits) into a fixed control-to-audio ring that process() drains once at the top of the block, so host parameter writes are race-free against the audio thread and land at the start of the next block (an overfull ring counts as a dropped event). The typed getters drain the ring first, so a set is immediately visible to a control-thread get.

  • YDSP codegen: the operand-register diagnostic skips scalar-load slot operands (operand 0 of loadParam/loadParamOut/loadStateF/loadStateI holds a state/param slot, not a value id), so state-heavy kernels no longer false-positive as having an undefined value.

  • YDSP optimizer: foldStateWriteBacks no longer folds a scalar state write-back whose produced value is read outside the move's own block - a function parameter reassigned inside an if keeps being read by the join block after the call, and folding the parameter copy renamed its producer and left that read with no definition at all, which asmjit rendered as <Reg-0>?255 on AArch64 (the ten YdspExamplePatchTests that failed that way are the end-to-end guard). The fold also stops treating a scalar load's slot index as a value id, requires matching value shapes, and applies overlapping write-backs into the same state register one at a time instead of as a batch against a stale snapshot. Dead-code elimination and the contraction pass now share one filtered value-use scan (forEachValueUse), so slot indices can no longer be mistaken for value uses anywhere.

  • YDSP optimizer: confirmed miscompiles fixed. Loop fusion retargets the fused loop onto the exit block its own loop record names (it used to aim the header one block before the loop, at the preheader), only drops a second loop's induction init when it really starts at the literal 0 (a 3..8 second loop no longer silently restarts at 0), and treats event emission and param/paramOut/event-field accesses as memory - two event-emitting loops no longer fuse. Constant folding no longer materialises a fold its own guard declined (x / 0, x % 0, x << 64 stay unfolded instead of becoming a compile-time 0), refuses to fold a non-finite or out-of-range ftoi, and rounds folded float32 values to float32 so folded and unfolded kernels agree. Strict math no longer applies the x * 0.0 -> 0.0 / x + 0.0 -> x identities (both are wrong for signed zero, NaN and infinity). Copy propagation no longer retargets a block's terminator condition past a redefinition of the moved register. A widened fmaF is refused by the scalar lowering instead of being scalarised per lane, a widened-reduction link must be single-use to be deleted, and dead event-field loads are now removed by dead-code elimination.

  • YDSP native codegen: the temporary YUP_YDSP_RA_DEBUG environment switch was removed (RA annotation is compiled out), a scalar-transcendental table miss now asserts instead of silently leaving the result register undefined, and the post-register-allocation redundant-move cleanup also removes identical fmov/ASIMD mov copies.

  • YDSP codegen (x86-64): integer variable-count shifts wrote the fixed physical cl register directly (mov cl, srcB.r8()). AsmJit's register allocator is blind to physical operands - it only models virtual ones - so it could keep another live value in RCX and the mov cl silently destroyed its low byte (most often a LICM-hoisted shift-count constant), corrupting the LFSR generator's w >> k / << 31 chain. The count is now passed as a virtual register and the allocator pins it into CL with full liveness knowledge.

  • YDSP codegen (x86-64): the shared clampI lowering selected the max into the result register and then ran the min with that same register as its own whenTrue operand. The x86-64 select lowering writes dst before reading whenTrue (mov + cmov), so the second stage degenerated to a no-op and clamp (n, -10, 10) always returned the upper bound - min/max/clamp/abs/sign on an int returned 6 where -14 was expected for negative n. The max stage now selects into a temporary, keeping the two select operands distinct on both backends.

  • Emscripten: PopupMenu's click-outside-to-dismiss and hover-forwarding (Desktop::handleGlobalMouseDown/Up/Move, fed from an SDL event watch in yup_Initialisation_sdl.cpp) built their MouseEvent from SDL_GetGlobalMouseState(), which reports real page-relative pixels on this platform. Component::localToScreen() instead anchors on SDLComponentNative::getPosition(), which returns (0,0) on Emscripten by design (there's no real window position to query, see the isMouseOutsideWindow fix below). The two frames only agreed when the canvas happened to sit near the page origin, so selecting a ComboBox/PopupMenu item would misfire "click is outside the popup" and dismiss it before the selection could register - reproducible on both HiDPI and standard displays, and sensitive to anything that shifts the canvas's on-page position (such as opening devtools). The global dispatcher now sources its position the same way regular per-window dispatch does on Emscripten, keeping desktop behavior (which needs true desktop-global coordinates) unchanged; mobile is left untouched pending separate verification.

  • Emscripten: isMouseOutsideWindow() compared SDL_GetGlobalMouseState() (real page-relative coordinates) against SDL_GetWindowPosition(), which is meaningless on this platform since SetWindowPosition is unimplemented in the SDL3 Emscripten backend and returns a value centered against a synthetic virtual display. This misclassified ordinary clicks inside the canvas as "outside the window" on every mouse-up, which called handleFocusChanged (false) and stopped SDL text input - silently dropping every typed character while raw key events (Tab, Backspace, Enter, and anything reading KeyPress directly, like MidiKeyboardComponent) kept working, the exact symptom reported against the new CodeEditor. The check is now skipped on Emscripten, where "mouse outside the window" isn't a coherent concept for a single always-present canvas.

  • YDSP codegen (AArch64): the register allocator was free to park a long-lived value in x30 (LR), which any call then destroyed. asmjit's AArch64 tables disagree with AAPCS64 about that register: a64func.cpp lists x30 among the callee-preserved GP registers, so the allocator believes a value there survives a call, while a64rapass.cpp only makes SP and FP unavailable - leaving x30 allocatable. blr overwrites LR with the return address, so any kernel that both calls out (a transcendental, or the exp behind smooth) and has enough simultaneously-live values to reach x30 had that value silently destroyed mid-body. The symptom depended entirely on what the lost value was doing: a state-array region base meant every later access through it addressed wild memory, which is how this was found - a fused kernel with four reverb-sized @ rings and a smoothed parameter faulting on str s26, [x30, x15, lsl #2] - but a promoted state scalar in x30 faults on nothing and simply makes the patch compute the wrong numbers, so this was a latent wrong-audio bug of unknown age rather than only a crash. x30 is now added to the function frame's unavailable set, which costs one of ~29 allocatable GP registers and is what a JIT should do with LR regardless; the fix sits on our side of the boundary so it survives an asmjit update. It needed three ingredients at once (a call, high register pressure, and a state array large enough for the corruption to land outside the allocation), which is why every existing patch and test missed it: four new YdspFusionTests shapes bisect exactly those ingredients, and the crashing combination is now a regression test.

  • YDSP runtime: three routing defects that only the old exactly-once rule kept unreachable, all fixed by the same new representation (a node's input wiring and the graph's output wiring are both CSR-flattened lists of connections, each owning its own delay ring). (a) outputSlotBuffer[s] was last-edge-wins, so a node output feeding both a graph output and a node input clobbered one marker against the other - and the two edge orderings failed differently. (b) outputSlotGraphOut[s] was a single int, so one output could not drive two graph outputs at all: the second silently replaced the first. (c) edge.delaySamples was read only on the node-input branch, so a.out -> [3] -> y was parsed, accepted and then silently dropped. Owning the delay ring per connection rather than per input slot is what makes x -> a.in; x -> [4] -> a.in; a comb filter instead of representable-only-once.

  • YDSP runtime: a connection { x -> y; } passthrough edge (graph input straight to graph output, whether hand-written or produced by inlining a pass-through subgraph) hit neither branch of the wiring producer, so it compiled and then output silence. It is now an ordinary graph-output source with no producing node.

  • YDSP: oversampling (node x = P * 4) on a non-float32 stream was silent memory corruption. rateMultiplier was validated only for the event-driven and subgraph cases - there was no type guard anywhere - and the runtime handed the node's buffers to Oversampler<float, ...>, so a float64 stream was read and written at half stride, past the end of the buffer. It is now a compile error naming the offending stream and its type. Inline delays already carried the equivalent guard.

  • YDSP: two out-of-bounds reads in the graph algebra's sequential composition. A graph input or output identifier used as a leaf gets arity (1, 1) but carries a port on one side only, so dry : wet produces a value claiming arity (1, 1) with no ports at all - the arity check then passes for dry : wet : Gain and the wire loop indexed outPorts[0] on an empty vector. Its mirror image, Gain : dry, indexed inPorts[0]. Separately, _ inside , contributes no ports while still adding to the arity, so (_ , Gain) - and (Gain , _) - claimed arity 2 with one port, and whatever was sequenced against it read inPorts[1]. The composition now verifies that a declared arity is backed by actual ports and reports which side has nothing to connect, and mixing _ with a port-carrying operand inside , is rejected with a message pointing at the connection form. (_ , _) still composes as it always did: it stays an identity.

  • YDSP runtime: prepare() appended to each node's inline-delay buffer list instead of rebuilding it, while the per-block path indexes that list 1:1 with the input slots. A second prepare() therefore left every slot pointing at the first call's ring, whose one-block scratch is sized for the old block size - so preparing at a larger block size wrote past the end of it. The list is now assigned to the input-slot count and filled by index, which removes the mismatch by construction. reset() also left the rings holding audio (the oversamplers beside them were already reset), so the pre-reset tail kept playing out for delaySamples more samples; they are cleared now. Both were uncovered because every existing test prepares exactly once. Relatedly, prepare() ran the one-shot init kernels before rebuilding the scratch arena and the node pointer tables; the init kernels are handed null stream pointers so nothing observable changed, but they now run last.

  • YDSP optimizer: algebraicSimplification treated any constant writing a register as that register's value, ignoring that the IR is not SSA and the register may be written again before the use being simplified. process { float t = 0.0; t = in * 2.0; out = t + 1.0; } folded t + 1.0 against t's declaration literal and emitted out = 1.0, dropping the input entirely - silent wrong audio, no diagnostic. Only a register defined exactly once is eligible now, which is the guard constantFolding already documents and depends on. The bug needed the reassignment and the later use to share a basic block, which is why the common float sum = 0.0; followed by an accumulating loop was never affected (a loop body is a separate block) - the kernel-fusion pass, whose junction locals are declared, assigned and read in one block, is what exposed it.

  • YDSP runtime: a connection with an inline delay (x -> [3] -> node.in) wrote the delayed samples back over its source buffer. For an edge fed by a graph input that is the caller's own array, which process() declares as Span<const float> and the runtime was const_cast-ing away: a host that reuses or shares an input buffer found it silently rewritten after every block, and a genuinely read-only buffer was undefined behaviour. The delayed samples now go into a per-edge scratch buffer allocated in prepare() and the node's input pointer is repointed at it, so the source is untouched; the sub-block dispatch paths re-derive from that pointer and needed no change. The feature had no test anywhere in the suite, which is why this survived - it now has three, including one asserting the input buffer is byte-identical after process(). The delay ring is float32-only (as oversampling already was), so an inline delay on a non-float32 stream is now a compile error instead of reading the buffer at the wrong element width.

  • YDSP codegen (x86-64): the backend did not compile at all. It told a float32 register apart from a float64 one with reg.typeId(), which the vendored AsmJit has no such member for - the API is snake_case throughout (is_vec128, new_gp32, virt_reg_by_reg) - so all thirteen call sites were hard errors, hidden on AArch64 hosts by the #if ASMJIT_ARCH_X86 guard around the file. The width was never available on the operand in the first place: unlike AArch64, x86 has no distinct register class per float width, so a f32 scalar, a f64 scalar and a packed f32 vector are all a 128-bit Xmm. It is read from the register's VirtReg instead, as the TypeId it was created with.

  • YDSP runtime: mid-block parameter automation on a polyphonic node reached voices 1..N-1 an entire sub-block early. A node's param block is shared by all its voices, but applyAutomation runs inside the per-voice loop and nothing rewound the block between voices, so voice 0 saw the correct pre/post values while every later voice entered the block with the param already holding voice 0's post-automation value - and therefore read the new value throughout its pre-automation sub-blocks. A Voice[4] bank with two notes held and a gain automated from 0.5 to 1.0 at sample 100 produced 1.5 before the offset instead of 1.0. The block-start value of every automated param is now snapshotted before the voice loop and restored at the top of each subsequent voice, so all voices replay the same timeline. Monophonic and single-voice nodes run the loop once and are untouched.

  • YDSP optimizer: a loop nested inside another loop overwrote the outer loop's recorded YdspLoopBound. The bound is resolved before the body is lowered but only stored on the YdspIrLoop afterwards, and it was held in a builder member in between, so both loops ended up reporting the inner bound and YdspKernelReport::boundedIterationCount under-reported the worst-case iteration count. The bound is now captured into a local at the point it is resolved.

  • YDSP runtime: a node's all-sound-off handling (MIDI CC120) kept a single sample offset per node per block, so two CC120s on different event inputs clobbered each other. The earlier (more urgent) silence point was lost, and a note re-triggered between two CC120s escaped the later one. Every CC120's offset is now recorded (fixed per-block capacity, reserved in prepare()) and applied as its own sub-block split point. Voice lookup is also scoped by event input now: each event input has its own MPEInstrument, so the same MIDI channel+note collides across inputs, and a note-off on one input could previously release a voice (poly) or remove a held note (mono) belonging to the other input.

  • YDSP compiler: syntax errors in an imported file were swallowed. The parser recovers from a syntax error via synchronize(), so the merged program was non-null and the import merged anyway, dropping the diagnostic; the merge now treats any recorded error like the top-level path does. The import rename pass now also covers processor-local function bodies, graph node parameter overrides and graph-leaf override expressions (plain-name library calls in those positions previously dangled after merging). The static recursion check now walks if/loop bodies and calls nested anywhere in an expression, so a recursive library function hidden in a conditional is rejected before inlining.

  • YDSP compiler: a diamond import silently dropped the second alias's namespace. The merge deduplicated by file path globally, so when two imported files both imported the same third file, only the first importer's namespace was populated - the second importer's references to its own alias of that file failed to resolve (and a same-file-different-alias import in one file lost the second alias too). The merge now deduplicates per (importing file, file, namespace) and runs each merge on a private copy of the parsed file, so every importer gets the shared declarations under its own namespace.

  • YDSP compiler: the intrinsic name list was hardcoded twice - a mirror of the semantic analyzer's table lived in the compiler's import rename pass, so the two could drift (and already had, once). The rename pass now queries the analyzer's table through a shared isIntrinsicName(). The function shadowing rule (a processor-local function beats a program-level one of the same name) was likewise duplicated between the semantic analyzer and the IR builder; both now resolve through one shared findFunctionInScope(). The module's runtime-initialised static lookup tables (intrinsics, builtin constants, IR op dispatch, annotation keys) are now constexpr arrays with no dynamic initialisation.

  • yup_core: new CountDownLatch synchronization primitive (addCount/countDown/wait), used by the parallel import parser to replace its hand-rolled mutex/condition-variable counting; the threaded import path now also has test coverage for diamond imports and failure cases.

  • YDSP runtime: sample-accurate sub-block dispatch advanced each sub-block's stream pointers from the previous sub-block's pointers while passing an absolute offset, so the offsets accumulated: any node split into three or more sub-blocks in one block (two or more distinct event/automation offsets) wrote its output - and read its input - at sum(offsets) instead of offset, corrupting samples past the end of the buffer in the worst case. Polyphonic outputs were immune (they re-derive from the per-voice scratch each call), which is why the existing voice-bank tests never caught it. Each sub-block now re-derives its pointers from the block's base.

  • yup_gui: mouse-event hit-testing now respects setWantsMouseEvents() across siblings, not just up the parent chain. New Component::findComponentAtForMouseEvent() walks the hierarchy the same way findComponentAt() does, but a component (or its whole subtree) that opted out of mouse events is skipped in favor of the next sibling underneath it, instead of only bubbling up to a parent; findComponentAt() itself is unchanged (still a pure, mouse-event-agnostic geometric hit-test). Previously an overlapping sibling that didn't want mouse events (e.g. a decorative Label drawn on top of a Slider) would still win hit-testing and swallow clicks meant for the sibling beneath it. SDLComponentNative::findComponentForMouseEvent() now delegates to findComponentAtForMouseEvent() instead of duplicating the ancestor-walk itself.

  • yup_core: Thread::getCurrentThreadId() is now TLS-based on wasm instead of returning the raw pthread_self() value, which is unreliable on the emscripten audio-worklet thread (a Wasm Worker, not a pthread) - thread-identity-based locks and checks (RecursiveSpinLock owners, ReadWriteLock writers, MessageManager thread checks) misbehaved there. Every thread (main thread, audio worklet, yup::Threads) gets a stable TLS id seeded by an atomic counter; yup::Threads sync their threadId to it at creation so getThreadId() keeps matching, while threadHandle stays the native pthread handle. The implementation must not consult getCurrentThread() - ThreadLocalValue uses getCurrentThreadId() to find its per-thread slot, which would recurse.

  • yup_audio_basics: AudioLockType (used by the MidiKeyboardState, Synthesiser and BufferingAudioSource locks) is a RecursiveSpinLock on wasm (never blocks - no futex_wait - and re-entrant, which Synthesiser requires when processing MIDI) and the re-entrant CriticalSection on desktop. MidiKeyboardState::allNotesOff additionally no longer re-enters the lock (allNotesOff/noteOff delegate to non-recursive *Locked helpers), so the state is safe even with a non-recursive lock.

  • yup_audio_devices: the locks taken on the audio thread - AudioDeviceManager::audioCallbackLock/midiCallbackLock (taken in every audioDeviceIOCallbackInt block), MidiMessageCollector::midiCallbackLock, AudioSourcePlayer::readLock and AudioTransportSource::callbackLock - now use AudioLockType instead of a raw CriticalSection, so on wasm a contended lock spins on the audio-worklet thread instead of blocking in a fatal futex_wait. getAudioCallbackLock()/getMidiCallbackLock() now return AudioLockType&.

  • YDSP wasm codegen: stream loads/stores (loadInput/loadOutput/storeOutput) treated the channel-pointer array (inputs/outputs) as the sample buffer - the channel pointer inputs[ch] was never dereferenced, so a kernel indexed &inputs[0] + i*4 and wrote sample data over the pointer array and into the surrounding heap. Every generated kernel that touched a stream corrupted the heap (crashes surfaced in the browser's audio worklet and in node tests). The channel pointer is now loaded first, then indexed by the sample.

  • examples/graphics: the "Pulse Bass" patch is now the demo's monophonic example - node voices = PulseVoice [[ mode: mono, priority: last ]] - and is the first patch to read e.isLegato: a fresh key press re-attacks at the new pitch immediately, while a legato continuation leaves the envelope running and glides the pitch over a new Glide parameter. The re-attack ramps from wherever the envelope currently is rather than resetting it to zero - with a single voice, resetting would chop a still-sounding release tail off mid-level and click, and ramping is how a mono synth's single-trigger envelope behaves anyway. Previously it was a 4-voice bank whose .isLegato flag nothing consumed, so the mono/legato path shipped untested by any example.

  • examples/graphics: the "YDSP Synths" patches clicked on every note. Poly Sine, Analog Saw, FM Bell and Wobble Lead all stepped their envelope straight to the note velocity in noteOn while the oscillator kept its previous phase, so the output jumped discontinuously at each attack (and at each voice steal). They now run a ~3 ms one-pole on the gain - the coefficient is a state float ... = 1.0 - exp (-samplePeriod / 0.003) initialiser, so it is computed once per voice in the init kernel rather than per sample. Pulse Bass already had a one-pole attack and is unchanged.

  • examples/graphics: the "Electric Piano" patch was ~24 dB quieter than the other synths (a * 0.125 output trim on top of partial weights that already sum to ~1); it now matches their level.

  • examples/graphics: the "YDSP Synths" demo only played patches with exactly one output stream - any other count fell into the silence branch, so a stereo patch compiled and ran but was never invoked at all (process() also rejects a buffer span whose size does not match the declared stream count). It now accepts mono and stereo patches: a mono graph is fanned out to every device channel as before, a stereo graph maps its two output streams to alternating channels, and the oscilloscope shows the mean of the two.

  • YDSP optimizer: a for loop nested inside an if hung the generated kernel. The if-lowering allocated its join block (and, with an else, the else block) before lowering the then region, so any block the region allocated landed past the join: the loop's preheader fell through into the join instead of into the loop header, and the join then fell through into that header - whose induction variable had never been initialised - so the loop ran forever. Both the join and the else block are now allocated after the regions they follow, keeping each region's blocks contiguous and the join last, which is the layout both the fallthrough-based asmjit codegen and the wasm backend's structured-region recovery require. The else if fix below only covered the else-carrying case.

  • YDSP optimizer: if / else if chains produced a non-linear block order - the outer join block was allocated before the inner if's blocks, so the outer join fell through into the inner then-block and the else-tail branch looped back into a cycle (infinite loop on the asmjit backend; "unexpected conditional branch" compile error on the wasm backend, which surfaced it first via the AnalogSaw patch's polyBlep). The join is now allocated after the else region, so nested else if blocks stay between the else block and the join, restoring the CFG-linear order both codegens rely on.

  • YDSP optimizer: loopInvariantCodeMotion treated loop-carried registers as invariant - the induction register is written both before the loop (sample mode: prologue constI 0; block mode: preheader movI) and inside the loop body, so its addI/movI update was hoisted out of the loop, freezing the counter and hanging the generated kernel (the sample-mode case slipped through even after the preheader exclusion because the prologue constI 0 is a constant op that re-seeded it). Only single-assignment values defined in the entry block (or constants) can now be invariant, registers redefined inside the loop (induction variables, path-dependent if/else values) are never hoisted, and loads from input/output streams, params and meters are not hoisted either (the loop body can overwrite that memory).

  • YDSP optimizer: constantFolding treated the non-SSA IR's value ids as single-assignment, so the sample-loop induction register - defined once by the prologue constI 0 and once per iteration by the movI update - folded its per-iteration update to a literal and the generated kernel looped forever (CompilesPassThroughAndProducesCorrectOutput hung). Constants are now only propagated for value ids defined exactly once.

  • YDSP optimizer: inlined func calls bound parameters directly to the caller's argument register (locals are mutable single-register IR slots), so a function that reassigns a parameter - e.g. t = t / dt in the AnalogSaw polyBlep - clobbered the caller's local and the block-end state store persisted the mutated value instead of the original (state[0] = t instead of phase), making the oscillator diverge to ±inf after the first cycle. Parameters are now pass-by-value: each argument is copied into a fresh value at the call site. Copy propagation also stops at redefinitions of the destination, so it can't propagate through the new parameter copies.

  • MidiKeyboardComponent: computer-keyboard input now actually matches keys (the previous full-KeyPress comparison never matched SDL events) and sends matching note-offs on keyUp, releasing all keyboard-triggered notes when focus is lost instead of leaving them stuck; the mouse wheel now scrolls left/right by single white keys (vertical wheel works on horizontal keyboards too) and Ctrl/Cmd + wheel zooms in/out, clamped between a single octave and the full 0-127 range.

  • YDSP optimizer: binary ops mixing a non-literal constant with a literal (e.g. int64(1) << 32) previously narrowed the constant to the literal's width, silently changing the result type (int64(1) << 32 computed as 32-bit and wrapping to 1); unifyOperands now mirrors the analyzer's contextual literal adaptation and only widens the source literal.

  • YDSP codegen: parameter and state loads now consistently read their slot index from the a IR field (was memIndex), fixing invalid memory offsets for params, prev/mem slots, @ delay write pointers and param-out meters.

  • YDSP optimizer: implicit int -> float coercion for literals and values used in float arithmetic, comparisons, intrinsics, and ternary/select branches; fixes invalid mixed integer/float instructions on AArch64 (fmul s0, s0, w0 and similar). sign() now converts its integer result to a float.

  • YDSP optimizer: the @ delay primitive delayed by n + 1 samples; the ring read now targets the slot one past the write pointer so x @ n delays by exactly n samples.

  • YDSP codegen: fmod(a, b) had no codegen case and silently produced garbage output; now lowered to a fmodf libm call.

  • YDSP codegen: integer / and % by a zero divisor now consistently return 0 on both x86-64 and AArch64 (previously trapped with SIGFPE on x86-64, and % returned the dividend unchanged on AArch64).

  • YDSP codegen: libm calls (sin, cos, tanh, pow, fmod, etc.) emitted an invalid blr with an immediate operand on AArch64; the call target is now materialized into a register first.

  • YDSP optimizer: mixing an integer literal with a float expression (e.g. 1 - a or 1 + 0.5 * side) lowered to integer arithmetic, coercing the float operands through fcvtzs and truncating fractional values; integer literals now adapt to the other operand's float width as the language's contextual literal adaptation requires.

  • YDSP semantic analyzer: writing 64-bit (float64/int64) input-value parameters in sample mode is now allowed (they double as per-sample accumulator registers); 32-bit parameters remain block-rate only.

  • YDSP semantic analyzer: a processor with two or more func definitions crashed during analysis - functionTable stored pointers into the growing functions vector, so registering the second function reallocated it and dangled the first function's entry (resolveFunctionCall then read freed memory). The vector is now pre-sized before registration.

  • YDSP optimizer: copyPropagation rewrote the a/b/c operand fields of every instruction, conflating value ids with the state/param slot indices carried by loadStateF/loadStateI/loadParam/loadParamOut (both live in the same small-int space). A sample-mode state reassignment redefines a low value id through a mov, so in a Reverb-style kernel (twelve @ delays interleaved with comb-state reassignments) the hidden ring write-pointer loads had their slot rewritten to a value id and read the pointer from inside the array segment instead of the scalar slot (load state+264 vs store state+108) - the ring store then indexed with garbage and crashed the generated kernel on the audio thread. Copy propagation now only rewrites operand fields that hold value ids for the target instruction.

  • yup_core (wasm): Time::getMillisecondCounterHiRes() - read every audio block by MidiMessageCollector::removeNextBlockOfMessages - called std::chrono::steady_clock::now(), which std::terminate()s (surfacing as an opaque abort()) when called from the browser's AudioWorkletGlobalScope: libc++'s steady_clock::now() goes through a pthread-aware clock_gettime path, and the audio-worklet thread is a lightweight Wasm Worker rather than a real pthread - the same "not actually a pthread" gap already worked around elsewhere for thread identity (Thread::getCurrentThreadId(), TLS-based on wasm) and for locking (AudioLockType, a non-blocking RecursiveSpinLock on wasm). getMillisecondCounterHiRes()/getHighResolutionTicks()/yup_millisecondsSinceStartup() now go through emscripten_get_now() on Emscripten instead - a thin performance.now() wrapper with no pthread machinery behind it, safe on every thread including the audio worklet. (An earlier attempt at this fix, changing the startup-time baseline from a function-local to a namespace-scope static, addressed a real but unrelated latent issue and did not fix this crash - the terminate was inside steady_clock::now() itself, not in its own fallback baseline.)

  • YDSP semantic analyzer: procStates (parallel to each state's per-processor index) was never cleared between processors, so a program with two or more processors declaring state could resolve a later processor's state against an earlier processor's YdspStateDecl - wrong type, wrong struct, silently wrong analysis. It is now reset alongside every other per-processor field.

  • YDSP optimizer: constantFolding computed divI/modI/shlI/shrI/addI/subI/mulI/negI/absI through plain int64_t arithmetic, which is undefined behaviour at the extremes (INT64_MIN / -1, a shift amount outside [0, 63], signed overflow near INT64_MIN/INT64_MAX) - triggered by the compiler's own fold, not generated code. divI/modI/shlI/shrI now leave the instruction unfolded (like the existing div-by-zero guard) when the operands would invoke UB; the others compute through uint64_t so the result matches the wraparound generated code would produce at runtime. YdspSemanticAnalyzer::tryConstantFold's shl/shr (used when folding a program-level let constant) had the identical unguarded shift and is fixed the same way.

  • YDSP runtime: getActiveVoiceCount() read a voice's [[ role: voiceActivity ]] flag through state.data() + offset unconditionally - state is only sized by prepare(), so calling it right after compile() (isValid() is already true then) read past the end of an empty buffer. It now falls back to the same conservative "treat as active" default used for a held voice.

  • YDSP runtime: an all-sound-off event (MIDI CC120) dropped for capacity (more than 16 in one block on the same node) still wiped the node's voice-slot bookkeeping unconditionally, freeing those voices for reallocation even though silenceVoice() never ran for the dropped offset - a note reusing the slot inherited the previous voice's filter/envelope state instead of a clean one. The bookkeeping is now only cleared when the event itself was actually recorded.

  • YDSP language: the recursive-descent parser, the semantic analyzer's AST walk, the IR builder's function-inlining chain and the wasm codegen's block/loop/if emission had no recursion-depth guard, so a pathologically nested expression, statement or call chain (not necessarily adversarial - a generated .ydsp source nests easily) could overflow the native stack instead of failing with a diagnostic. All four now share one YdspRecursionGuard depth limit.

  • YDSP language: the lexer narrowed each source character to unsigned char in peek()/current(), so a non-ASCII character whose low byte matched a whitespace/newline byte could end a // comment early (leaking the rest of the comment as real tokens) or desync line/column tracking; string-literal content was separately narrowed to char, corrupting non-ASCII text. Both now carry the full code point.

  • YDSP language: parseProcess()/parseEventHandler() pushed parseStatement()'s result into the process/event-handler body unguarded, unlike the otherwise-identical pattern in parseBlockStatement()/parseFunction() - a malformed statement (e.g. a bare ;) landed as a null entry in that vector. YdspParser::parseProgram()'s doc comment also claimed it returns nullptr on any syntax error, which it never does (parsing recovers and returns a non-null, structurally complete program - callers must check diagnostics.hasErrors()); the comment now matches the actual, intentional contract.

  • YDSP runtime: droppedEventCount was a plain uint64_t incremented on the audio thread and read from a UI/control thread with no synchronization - now a relaxed std::atomic. Several byte-buffer reads/writes (constant folding into state, parameter automation snapshot/restore, get/setParameter, meter reads) aliased a uint8_t* through a differently-typed pointer, which is undefined behaviour when the buffer offset isn't aligned for the target type; all now go through memcpy, matching a precedent already used elsewhere in the same file.

  • YDSP backend: YdspCompiledKernel's wasm module lookup (wasmModules/wasmIndex) had no bounds check, so a kernel wrapper that outlived a graph rebuild could index past the end of the module vector; it's now bounds-checked before every access. Function-pointer conversions were scattered across several ad-hoc reinterpret_casts (contradicting a comment claiming they were "confined" to one class) and are now centralized behind two small helpers (ydspFnPtrCast, ydspFnPtrToInt64) in yup_YdspAbi.h. The wasm text decoder's LEB128 reader could shift a 32/64-bit accumulator by more than its width on a malformed byte sequence (UB); it now caps the shift.

  • YDSP backend (emscripten): a pointer was narrowed to the wasm-call context id via a direct reinterpret_cast<int>, which isn't a well-defined conversion in general (it happens to work only because this file is wasm32-only, where pointers and int are both 32 bits) - now goes through uintptr_t to make that dependency explicit. yupDspWasmFreeKernel also skipped the same registry-exists guard every other function in the file uses, so freeing a kernel in a realm that had never registered/called one threw instead of no-oping.

  • YDSP codegen (AArch64): scalar float32 values were allocated through the packed kFloat32x1 TypeId, which AsmJit maps to a 64-bit (D) register on AArch64 but a 128-bit XMM on x86-64. Every float32 value was then computed at 64-bit width, loads over-read the next slot, and scaled array/stream stores (str d0, [base, index, lsl 2]) failed to assemble with InvalidAddressScale, breaking AnalogSaw on arm64. The scalar TypeIds are used again on AArch64; x86-64 keeps the packed forms.

  • yup_audio_gui: MidiKeyboardComponent repainted synchronously from its MidiKeyboardState::Listener callbacks, but those callbacks are documented to be able to fire from an audio or MIDI input thread - feeding the state via processNextMidiEvent() from a CoreMIDI callback repainted off the UI thread, racing the paint loop (stuck key highlights) and crashing in Component::repaint. The listener callbacks now only trigger an internal AsyncUpdater; the repaint is coalesced and applied on the message thread, so the key state may be updated from any thread. Regression-tested in yup_MidiKeyboardComponent.cpp by feeding notes from a worker thread and asserting the deferred repaint lands on the message thread.

  • Examples: the YDSP Synth Lab demo (examples/graphics/source/examples/YdspSynths.h) registered its MidiMessageCollector directly as the MidiInputCallback, so hardware MIDI went straight into the audio queue and never reached MidiKeyboardState - incoming notes never highlighted the on-screen keyboard, and incoming pitch bend / CC1 never moved the wheels. The demo is now the callback itself: it feeds keyboardState (which forwards note on/off to the collector through its existing keyboard-state listener), maps incoming pitch bend and CC1 onto PitchWheelComponent/ModWheelComponent with dontSendNotification (applied in refreshDisplay() on the message thread), and forwards every other message to the collector.

  • YDSP wasm backend: three defects fixed. The vectoriser widens float compare/select chains by reusing the scalar eqF..geF/selectB opcodes on lanes > 1 values, and the wasm codegen lowered every compare to a scalar f32.gt/f32.eq over v128 operands (the only branch was 32/64-bit) and a widened select to the untyped select - so any vectorized kernel containing a comparison produced a module the browser refused to compile (expected type f32, found local.get of type v128, the YdspGraphTests/YdspVectorizerTests compare shapes). The codegen now emits the f32x4.eq..ge per-lane mask opcodes and lowers a widened select through v128.bitselect, matching the AsmJit cmpps/fcmeq plus andps/bsl semantics. Second, wasm ignored fastMath entirely (it was clamped false), so a * b + c stayed two separate float32 roundings - TheContractedSubtractRoundsOnce produced 0.789999962 instead of the fused 0.790000021. wasm now honors fastMath like native: scalar float32 contraction is fused there too, expanded through the exact float64 sequence (the target has no fused multiply-add instruction), so it rounds once and stays bit-stable with the native default. Because a widened chain cannot fuse without a packed fused instruction, the vectoriser keeps the implicit per-sample stream loop (runtime blockSize bound) scalar when it holds a fusable mul->add/sub chain on such a target (new rejection reason keptScalarForContraction) so the chain still rounds once; constant-bound bank for i in 0..N loops keep widening unfused by design. Third, every successful wasm kernel compile rendered its whole module to text as an info diagnostic, accumulated through yup::String += on a String with no capacity reserve - quadratic in module size, so the largest example patch (TX81Z) effectively never finished compiling on wasm; the renderer now accumulates into a growth-friendly std::string.

DSP

  • yup_dsp_jit: YDSP gains parameter smoothing. The new smooth (x, tau) intrinsic is a one-pole ramp towards x with a tau-second time constant, plus a snap on the first sample where the ramp cannot advance, so the target arrives exactly and leaves no denormal tail (the compare tests the step rather than the remaining distance: a float32 lerp stalls while still short of the target by roughly ulp / coeff, so any fixed epsilon would either be unreachable or truncate a fast ramp); it lives under the same restrictions as the delay primitives (per-sample body, outside loops, not in event handlers) and costs one hidden float slot plus one hidden int slot per call site (the coefficient's exp is loop-invariant and hoisted out of the sample loop). [[ smoothing: <seconds> ]] on an input value float is sugar for it: one synthetic local is prepended to the per-sample body and that body's references to the endpoint are rewritten, so event handlers, func bodies and getParameter() keep seeing the raw target. Nothing changed in the IR ops, the codegen backends or the runtime - automation is still a step, the ramp happens in generated code. Every "YDSP Synths" example patch now smooths the parameters that actually step the signal - Delay feedback/damping/mix, Chorus depth/mix, Compressor threshold/ratio/makeup, Distortion drive/tone/mix, Reverb mix/damping/room size, Analog Saw resonance, FM Bell mod index, Wobble Lead both cutoff extremes, Pulse Bass width and Electric Piano's tremolo depth - so dragging a knob no longer zippers. Parameters that only set a rate (LFO rates, the compressor's attack/release) or that are read once in an event handler (envelope times, FM ratio, all the Electric Piano voice controls) are deliberately left stepped. Delay's time stays raw on a slew-rate argument: ramping an ~88000-sample read index over 20 ms scrubs the buffer at ~100 samples per sample, worse than the single jump it replaces - whereas Chorus's depth spans only ~880 samples and so slews at roughly the rate its own LFO already does. Analog Saw and Pulse Bass smooth their cutoff at the filter coefficient rather than at the parameter - float k = smooth (1 - exp (...), 0.02) - because a parameter is sampled once per kernel invocation, so an expression built only from parameters is already hoisted out of the sample loop; smoothing the parameter itself would pull that exp back in, while smoothing the coefficient keeps it hoisted and costs one lerp. The four voices' hand-rolled envSmooth/smoothCoeff anti-click one-pole collapses to a single smooth (env, 0.003); not quite a pure refactor, as a voice's very first note now attacks slightly faster (the primed flag snaps the gain on the voice's first sample instead of ramping it up from zero). That is inaudible - the oscillator and filter state are zero at that point too, so the voice still starts from silence - and a stolen voice is unaffected, since state is not zeroed on steal and the smoother ramps exactly as before.
  • yup_dsp_jit: YDSP gains four small language additions, all const-folded away before name resolution so nothing downstream of the semantic analyzer changes. Program-scope let name = <constant expression>; declares a compile-time constant usable as a state array size, a for bound or in any expression (imported constants are namespaced like imported processors, and redeclaring a constant's name is an error); state declarations take initialisers (state float feedback = 0.5;, state float table[32] = { ... }; with trailing elements left zero) which are lowered into the processor's init kernel, synthesising one if the processor has no explicit init { } block; the new samplePeriod builtin is 1 / sampleRate (loop-invariant, so the division is hoisted out of the sample loop); and [[ init: <value> ]] is accepted as an alias for a parameter's default value, so annotation blocks paste in unchanged (an explicit = <value> still wins). Constant folding now also handles binary arithmetic, so an endpoint default like = 1.0 / 3.0 evaluates instead of silently becoming zero.
  • yup_dsp_jit: YDSP for loop variables are now scoped to the loop body as documented - sibling loops may reuse the conventional for i without colliding, and the variable is no longer visible after the loop (previously the second for i in one body was rejected as a duplicate symbol).
  • examples/graphics: new "Electric Piano" YDSP synth patch (data/synths/ElectricPiano.ydsp), a 16-voice additive electric piano - a 32-partial complex-rotation oscillator bank per voice with velocity-blended spectra, per-partial decay interpolated in 64-sample chunks, and a stereo triangle tremolo. The bank is stored as parallel state float[32] arrays rather than a struct of complex values so the harmonic loop walks unit-stride. The two partial-weight spectra are hand-authored, not measured from a real instrument.
  • yup_dsp_jit: YDSP is now a full MIDI/MPE instrument host. Processor-scope events grew from noteOn/noteOff to seven shapes named noteOn (.pitch, .velocity, .isLegato), noteOff, pitchBend (.bendSemitones), pressure, slide, control (.control, .value) and programChange (.program) - driven by a single shape/field table that the analyzer, the IR builder and the runtime all consult, and lowered through two offset-carrying IR ops (loadEventFieldF/loadEventFieldI) that replace the previous per-field opcodes. Ingestion now goes through yup::MPEInstrument, so plain MIDI and MPE share one path and note expression is keyed by note identity rather than by pitch: per-note bend/pressure/slide reach only the owning voice, sustain, sostenuto, reset-all-controllers and all-notes-off are honoured as ordinary note lifecycle, and all-sound-off (CC120) silences and re-runs init at its exact sample offset. Voice behaviour is declared on the node - node v = Voice[8] [[ mode: poly, stealing: oldest ]], node b = Bass [[ mode: mono, priority: last ]] - with mono mode keeping an allocation-free held-note stack and flagging legato continuations through .isLegato. New YdspAudioGraph::setMpeZoneLayout() / setLegacyMidiMode() (legacy is the default, so existing hosts are unaffected), and prepare()'s per-voice event capacity default rose from 32 to 64 for MPE traffic. Two behaviour changes: a second note-on at the same pitch and channel now retriggers (releases) the first instead of stacking a second voice, and expression for a note with no allocated voice is discarded and counted in getDroppedEventCount().
  • yup_dsp_jit: the native asmjit backend was split per architecture - YdspCodegen is now YdspAsmJitCodegen (architecture-independent facade) with the shared lowering in YdspAsmJitCodegenImpl and the per-target encodings in YdspAsmJitCodegenX64 (x86-64 SSE) and YdspAsmJitCodegenARM64 (AArch64 ASIMD) (backend/yup_YdspAsmJitCodegen{,.X64,.ARM64}.*); purely organizational, no behavior change beyond the class rename.
  • yup_dsp_jit: new WebAssembly backend for emscripten/browser targets (previously the module hard-errored on YUP_WASM). Each kernel is lowered to a self-contained wasm module (backend/yup_YdspWasmEmitter.* binary writer + yup_YdspWasmCodegen), run by the browser's native WebAssembly engine: the module imports the host's shared linear memory as env.memory (host buffers addressed directly as i32 offsets, no marshalling), libm intrinsics come from env backed by Math.* with C-exact round/copysign/fmod, and structured control flow lowers to wasm block/loop/if (zero-division-guarded divI/modI, trunc-based modF). Kernels are instantiated per JS realm (main thread and audio-worklet thread each get their own copy, lazily on first use) and are keyed by unique per-kernel ids, so a worklet realm can never invoke an older graph's kernel with a newer graph's context when patches are swapped; YdspAudioGraph::prewarmKernels() pre-registers them in the calling realm, and the graphics example prewarms its graph from the audio callback. The asmjit dependency is now desktop-only, the wasm tests run under the emscripten node target, and the examples/graphics YDSP Synths demo works in the browser - its "Dump Asm" action prints the generated wasm as WebAssembly text (YdspWasmCodegen::toText, recorded as an info diagnostic at compile time, mirroring the asmjit assembly log on desktop). Defined-function indices are assigned in the wasm function index space (imported functions only; the memory import no longer shifts them).
  • yup_dsp_jit: the whole module now uses yup::String / yup::StringRef instead of std::string / std::string_view - AST fields, tokens, parser/analyzer/optimizer/codegen signatures and diagnostics messages - and reuses yup::String facilities (getLargeIntValue(), getDoubleValue(), substring(), lastIndexOfChar(), concatenation) in place of the stdlib string helpers.
  • yup_dsp_jit: node-level oversampling (node = P * N) now resamples through yup::Oversampler (windowed-sinc, 2x/4x/8x) instead of the hand-rolled linear-interpolation upsample and boxcar-average downsample; the resampler introduces 2 * SincRadius (16) samples of latency per oversampled node, oversampling requires equal input/output stream counts (or no inputs), and unsupported factors run the node at 1x.
  • yup_dsp_jit: YdspAudioGraph now exposes host-UI parameter metadata - getParameterCount()/getParameterInfo(slot) (qualified name, [[ name ]] display name, type, declared default, [[ min ]]/[[ max ]] bounds) plus getInputStreamCount()/getOutputStreamCount() - so hosts can build sliders directly from a patch. New examples/graphics "YDSP Synths" demo (desktop-only, since yup_dsp_jit requires asmjit) loads .ydsp synth patches from data/synths/, compiles them lazily with YdspCompiler, builds parameter sliders from the annotations, and drives them from a MIDI keyboard through YdspAudioGraph::process(..., MidiBuffer*, ...).
  • yup_dsp_jit: YDSP gains MIDI-driven events - input event noteOn/noteOff endpoints, event <endpoint> (<param>) { ... } handler blocks, and node = Processor[N] voice banks with fixed-size, allocation-free voice stealing. MIDI byte-decoding and voice allocation live in the runtime (YdspAudioGraph); events and parameter automation are dispatched sample-accurately via runtime sub-block splitting (the same compiled kernel is re-invoked per sub-block - no codegen changes), which also powers a new stepped parameter-automation API (YdspAutomationEvent, YdspAudioGraph::getParameterSlot()); YdspAudioGraph::process gains a MidiBuffer/automation overload and getDroppedEventCount(). A graph-level input event midi; documents MIDI consumption; routing is structural (broadcast to every event-driven node). Fixed a pre-existing out-of-bounds param/meter indexing bug in validateConnectivity for stream-free, parameter-carrying processors.
  • yup_dsp_jit: YDSP processors can now declare struct types with primitive and fixed-array fields, and state variables can be struct instances (Voice state;) or arrays of structs (Voice voices[32];). Fields are accessed with . (state.phase, voices[v].gate, state.buf[i], voices[v].buf[j]) and flatten into the existing scalar/array state slots - no IR or codegen changes required. Struct states are state-layout only (no struct values/locals/params yet).
  • yup_dsp_jit: YDSP processors can now declare an init { ... } block that runs once before audio starts (state writes, param reads, bounded loops; streams and delay primitives are rejected). Each init block compiles to a one-shot kernel sharing the processor's state layout; YdspAudioGraph::prepare() runs them in topological order, and the new YdspAudioGraph::reset() re-zeroes state and re-runs them.
  • yup_dsp_jit: YDSP now supports the C-style bitwise operators &, |, ^, <<, >>, ~ (and their compound assignment forms &=, |=, ^=, <<=, >>=) on int32/int64 operands, with C-compatible precedence (| < ^ < & < equality < relational < shift < additive), arithmetic (sign-propagating) right shift, and full support through the optimizer (constant folding + identity peepholes) and the asmjit backend on both x86-64 and AArch64.
  • yup_dsp_jit: YDSP now supports int32/int64 and float32/float64 primitive types end-to-end (float/int are aliases for the 32-bit types). Typing is strict: no implicit conversions, contextual literal adaptation only, with explicit int32(x)/int64(x)/float32(x)/float64(x) casts; if/&&/||/! require bool, indices and loop bounds require int32, and the delay primitives stay float32.
  • yup_dsp_jit: float64/int64 streams, parameters and meters are supported through the whole pipeline: the asmjit backend emits SSE2 double instructions (movsd/addsd/comisd/cvtsd2ss/…) and 64-bit GPR arithmetic on x86-64, register-sized fadd d/fcvt/sxtw/scvtf on AArch64, and width-matched libm variants (sin vs sinf). YdspAudioGraph::process gained per-stream/param element-type introspection and typed parameter accessors (getDoubleParameter/setIntParameter/…).
  • yup_dsp_jit: internal module sources were reorganized for maintainability - the semantic analyzer, the IR optimizer and the graph runtime monoliths were split into single-responsibility files (analysis/semantic/, optimiser/builder/ + optimiser/passes/, runtime/graph/); purely organizational, no API or behavior change.
  • examples/graphics: the "Analog Saw" YDSP synth patch now uses a polyBLEP band-limited sawtooth (naive ramp minus the polyBLEP residual at the phase wrap) instead of the naive aliasing ramp, demonstrated through the polyBlep function.
  • examples/graphics: every oscillator left in the YDSP synth patches under data/synths/ is now band-limited too - the remaining naive saws (Haas Widener, Parallel Drive, Parallel Rack, Wobble Lead) and the Pulse Bass pulse use the same polyBLEP treatment (two edges for the pulse), and the FM Bell is rebuilt as an additive (Le Brun) expansion whose Bessel-weighted sidebands are silenced above Nyquist, so none of the demo voices alias.
  • examples/graphics: new "Wave Lab" YDSP synth patch (data/synths/WaveLab.ydsp) demonstrates four polyBLEP band-limited oscillators behind one wave selector knob - saw, square, triangle (whose corners are rounded by the exact antiderivative of the BLEP residual) and a width-morphable pulse.
  • examples/graphics: the "YDSP Synths" demo gains a "Dump Asm" button that prints the asmjit assembly listing of the current patch's compiled kernels to the console (replayed from the graph's compile-time diagnostics).
  • examples/graphics: the "YDSP Synths" demo gains Performance/Editor tabs - Performance keeps the existing knob grid, keyboard and oscilloscope; Editor is a live CodeEditor (YDSP syntax highlighting) bound to the selected patch's source with a Compile button that recompiles it through YdspCompiler and swaps in the resulting YdspAudioGraph on success (refreshing the Performance tab's knobs), or shows the compiler's diagnostics (source line + caret) in a red panel on failure.

Audio

  • yup_audio_basics: new AudioLockType alias - yup::RecursiveSpinLock on wasm and yup::CriticalSection elsewhere - used by the module's audio types (MidiKeyboardState, UMPKeyboardState, Synthesiser, MPESynthesiser*, MPEInstrument, and the audio sources). On wasm a CriticalSection can block in a futex_wait, which is fatal on the browser audio-worklet thread; RecursiveSpinLock busy-waits on an atomic and never blocks, while keeping the re-entrancy of CriticalSection.
  • yup_audio_basics: MPEInstrument can now be driven from an audio callback without allocating - new reserveNotes (int) preallocates the note-tracking storage, and note removal uses the Array::remove() (in yup_core) which is now not allocating on removal.

Thirdparty

  • New asmjit_library module (thirdparty/asmjit_library): AsmJit machine-code generation library (core, x86, AArch64 and ujit backends), statically linked via the YUP module system.

Breaking changes

  • Component transforms now apply hierarchically: a parent's transform carries its children when painting, hit-testing, delivering mouse and drag-and-drop events, and in localToScreen() / screenToLocal() and the helpers built on them. Inside paint(), Graphics::getTransform() holds the linear part of the composed transform and the drawing area position the translation, so untransformed hierarchies paint exactly as before. The component transform is no longer applied twice to effect and cached composites.
  • yup_dsp_jit: the public API now uses yup::String / yup::StringRef instead of std::string / std::string_view - YdspCompiler::compile, every YdspAudioGraph param/meter accessor, YdspDiagnostics::addError/addWarning/addInfo/setSource/toString, YdspDiagnostic::message, YdspKernelReport::name/loopBounds (now yup::StringArray) and YdspParameterInfo::name/displayName. Callers passing const char* or std::string keep working unchanged (implicit StringRef conversion); explicit std::string_view arguments must switch to yup::StringRef.
  • yup_dsp_jit: YdspAudioGraph::process() no longer takes raw void* stream pointer tables. Stream buffers are now passed as typed spans - process (yup::Span<const YdspInputBuffer>, yup::Span<YdspOutputBuffer>, int) - where YdspInputBuffer/YdspOutputBuffer are std::variants of yup::Span<float|double|int32_t|int64_t>. The active variant alternative carries the buffer's element type, so a mismatched buffer is ignored and reported through the returned YdspProcessResult value instead of being reinterpreted; the generated-kernel ABI (YdspKernelContext) is unchanged. The process32() convenience overload was removed - callers pass typed spans to process() directly.
  • macOS: OpenGL rendering backend removed in favor of Metal only
  • LottieReader::parseFile(), parseData(), parseStream(), and parseFromZip() now return ResultValue<AnimationComposition::Ptr> and no longer take a trailing String* outError out-parameter; check wasOk()/failed() and read the message via getErrorMessage().
  • AnimationFrameExporter is now an instance-based class bound to a GraphicsContext (construct AnimationFrameExporter exporter (ctx); then call exporter.renderFrame(anim, …) / exporter.renderAllFrames(…) / exporter.exportToGif(anim, …)), so it can own and reuse the GPU matte-composite pipeline across frames instead of recompiling it per frame. The exportToGif(frames, frameRate, …) frame-sequence encoder remains a static helper.
  • Config macro YUP_EMBED_DEFAULT_THEME_TEXT_FONT renamed to YUP_EMBED_DEFAULT_THEME_TEXT_SERIF_FONT; the embedded default text font now only covers the serif font. A new YUP_EMBED_DEFAULT_THEME_TEXT_MONOSPACE_FONT config selects whether the monospace theme font is embedded.
  • SyntaxDefinition no longer carries colors: getColor(), getSelectionColor() and the JSON "colors" section were removed. Token and editor colors now live in the new CodeEditorScheme (see CodeEditor::setScheme).
  • Artboard no longer defines its own Layout and Alignment enums; it uses YUP's shared Fitting and Justification types. setLayout/getLayout became setFitting/getFitting taking a std::optional<Fitting>, and setAlignment/getAlignment became setJustification/getJustification taking a Justification. The mapping is fill→Fitting::fill, contain→Fitting::scaleToFit, cover→Fitting::scaleToFill, fitWidth/fitHeight unchanged, none→Fitting::none, scaleDown→Fitting::centerInside, and layout→std::nullopt (no fitting of our own, so the artboard resolves its authored Rive layout constraints). Alignment maps 1:1 onto the Justification corner/edge names, with topCenter, bottomCenter spelled centerTop, centerBottom. Fitting values Rive has no equivalent for (tile, centerCrop, stretchWidth, stretchHeight) are accepted and fall back to scaleToFit; because Justification is a bitfield it also admits partial states the old enum could not express, and an axis with no flag set is centered on that axis. NodeAttachmentOptions::pivot/anchor already took a Justification, so the class now speaks one alignment vocabulary throughout.
  • Font loading is now static-only: the instance loadFromData() / loadFromFile() methods were removed in favor of Font::loadFontFromData(), Font::loadFontFromFile(), Font::loadFontFromFirstAvailableFile(), Font::loadSerifSystemTextFont() and Font::loadMonospaceSystemTextFont(), all returning ResultValue<Font>.
  • The yup_rhi descriptors now own their data instead of pointing at caller-managed storage. GpuVertexBufferLayout::attributes and GpuPipelineOptions::vertexBuffers are std::vector (dropping attributeCount and vertexBufferCount), GpuPipelineOptions::colorTargets is a std::vector capped at four (dropping colorTargetCount), GpuShaderSource's code / bindingMap / glFixup are now std::vector<uint8> (owning) instead of Span<const uint8> with entryPoint a String, and GpuTextureDesc::label / GpuSamplerDesc::label became String. Call sites that built a static constexpr attribute table to keep it alive can now build the layout inline; a options.colorTargets[0].format = … becomes options.colorTargets.emplace_back().format = …. Use gpuShaderSourceBytes() to build a GpuShaderSource::code blob from source text. GpuDevice::beginOffscreen now takes the new GpuFrameDescriptor instead of rive::gpu::RenderContext::FrameDescriptor.
  • GpuDevice's move constructor and move assignment are now = delete. They were publicly defaulted on a ReferenceCountedObject, so moving a device out from under live GpuDevice::Ptr holders corrupted the refcount.
  • GpuBuffer::Impl, GpuBuffer::getImpl() and GpuBuffer::createWithImpl() moved from public: to private: (they were marked @internal by comment only); the backend factories that use them are friends. GpuTexture::getOreTexture() and GpuTexture::getRenderImage() were removed - neither had any caller.
  • DragAndDropData is now a MIME store and moved from component/ to the new dragdrop/ folder. Payloads are held as an Array<ClipboardData> under well-known MIME types (DragAndDropData::mimeTypeText / mimeTypeUriList / mimeTypePng), with text, files, URIs and images exposed as convenience accessors over that single store plus an optional same-process var native object. Consequently getFiles(), getText() and getUris() now return by value (Array<File> / String / StringArray) rather than const& (they decode from the MIME store), and the class gained withImage / withMimeData / withNativeObject with the matching get* / has* / getMimeTypes / getAllMimeData accessors. There are no lazy or promised data providers: every MIME blob is an eagerly-owned MemoryBlock.
  • The five drag-and-drop virtuals on Component (isInterestedInDrag, itemsDropped, itemDragEnter, itemDragMove, itemDragExit) and the Component::internalItemDrag* dispatch they fed have been removed, so Component no longer carries any drag-and-drop surface. Drop targets are now an opt-in mixin: derive from DragAndDropTarget (in dragdrop/) alongside Component and override isInterestedInDragSource / itemDropped / itemDragEnter / itemDragMove / itemDragExit — or assign the matching std::function members — each receiving a single DragAndDropSourceDetails that carries the payload, the source component, the target-local position, the allowed actions and the suggested action. The library finds targets with a dynamic_cast<DragAndDropTarget*> (see DragAndDropTarget::dispatchItemDrop() and friends), preserving the previous enter/move/exit and child-to-parent drop-bubbling semantics. The Python bindings for the removed Component hooks were dropped; DragAndDropTarget and DragAndDropTargetComponent are bound, as are DragAndDropSource and its DragOptions.
  • Oversampler was renamed SincOversampler (resampling/yup_SincOversampler.h) now that it is one of two oversampler designs, and the Oversampler2xFloat … Oversampler8xDouble aliases were removed: HalfbandOversampler2xFloat … HalfbandOversampler32xDouble name HalfbandOversampler instantiations with the default design (100 dB, passband to 0.45 of the input rate, stopband from Nyquist, linear-phase FIR). The call surface is identical, but getLatencyInSamples() reports the cascade's own delay (about 140 input samples for the 4× round trip with the defaults, against 32 for the radius-16 sinc) and getGenerationLatencyInSamples() is no longer static constexpr. Code that needs the sinc design or a non power-of-two factor should spell out SincOversampler<SampleType, Factor, Radius>
  • ComponentNative::setFocusedComponent() takes a second FocusChangeType argument saying what moved the focus. It defaults to FocusChangeType::focusChangedDirectly, so callers are unaffected, but any class implementing the pure virtual has to match the new signature.

Core

  • Added File::bundleDirectory and File::hostBundleDirectory special locations: the resource root bundled with the current YUP binary itself (a plugin's own .vst3/.component/.clap, an app's own bundle, or an Android APK's assets) versus the host application's bundle when running as a plugin. On Android, files under bundleDirectory are read directly out of the APK via AAssetManager through the normal File/FileInputStream API - no first-run copy to disk - and are read-only (createOutputStream(), deleteFile(), createDirectory() fail cleanly). Files can be placed there with the new yup_add_bundled_resources() CMake function
  • Added a YAML class: a self-contained YAML parser and writer converting between YAML text and var (parse, fromString, toString, writeToStream, FormatOptions), with core-schema type resolution, block/flow collections, block scalars, and anchors/aliases/merge keys
  • Added a CancelToken class (threads/yup_CancelToken.h): a thread-safe, copyable observer token with wasCancelled(), blocking observation via waitForCancellation(), and callback observation via registerCallback()/Registration
  • Added a CancelTokenSource class (threads/yup_CancelTokenSource.h): a move-only RAII owner of a CancelToken that is the sole canceller, requesting cancellation automatically when destroyed (unless moved-from), with observer copies obtained via getToken()
  • Fixed IPAddress disagreeing with itself about byte order inside each 16-bit group. The class reinterprets its byte array as uint16 groups on little-endian hosts, but the string parser packed a parsed group the opposite way round, so an address parsed from "fe80::1" did not equal the same address built from its integer groups, and an IPv4-mapped address did not round trip through toString() and back or compare equal to its IPv4 counterpart. toString(), both uint16 constructors, the parser and convertIPv4MappedAddressToIPv4() now pack and unpack a group through the same low-byte-first layout
  • WaitableTimer now meets its deadline on Apple platforms with a short series of absolute mach_wait_until() sleeps, each covering 75% of the time still left, followed by a 50us busy-wait. Darwin coalesces timer expiries into a window proportional to the requested sleep duration, so the previous condition-variable wait overshot in proportion to how long it was asked to wait, whereas shrinking sleeps make that error decay geometrically over a handful of syscalls. The condition-variable fallback is no longer compiled on Apple
  • WaitableTimer runs the same cascade on Linux and Android using absolute clock_nanosleep() sleeps, and asks for the smallest possible timer slack on the calling thread the first time it waits (Linux applies 50us of slack to every timer expiry outside the realtime scheduling classes, which an application is not normally allowed to join). The busy-wait budget is 150us there rather than Apple's 50us, measured on a Galaxy S25 running Android 16: 50us leaves a p99 wake error of 180us, 150us brings it to 86us, and going higher only costs more CPU. The condition-variable fallback now only remains for Windows (when the waitable timer cannot be armed) and WebAssembly
  • Fixed WaitableTimer::waitUntil() returning immediately for every deadline within one millisecond. The platform-specific final wait now runs for short deadlines instead of starting time-sensitive work early
  • Fixed Thread::startRealtimeThread() leaving the thread unstartable on macOS/iOS when the realtime upgrade failed: createNativeThread() reported the failure but kept the handle of the thread that had already bailed out, so isThreadRunning() stayed true and any later startThread() was refused

Audio

  • Added HalfbandOversampler (yup_dsp/resampling/), a power-of-two oversampler built from a cascade of 2× halfband stages so that only the stage next to the base rate has to be steep. Design selects the family and targets: linearPhaseFIR designs Kaiser-windowed halfbands, verifies every stage's stopband numerically at prepare() and keeps both latencies whole input samples by rounding the cascade's delay at the top rate; polyphaseIIR designs elliptic halfbands as two allpass branches (Valenzuela & Constantinides) for a few multiplies per sample and minimal, frequency-dependent latency. Defaults are 100 dB of rejection, a passband to 0.45 of the input rate and, for the FIR, a stopband starting at the input Nyquist, so nothing folds back into the band (a pure halfband first stage, selected with stopbandEdge = 1 - passbandEdge, costs a quarter as much but folds 0.5 to 0.55 of the input rate into the top of the band, which is what the IIR family always does). With the defaults 4× decimates in about 300 MACs per input sample against 128 for the radius-16 SincOversampler at 90 dB, a 0.36 passband and a 40 dB leak at Nyquist; at 32× it is about 600 against 1024, and the IIR cascade needs about 76 multiplies. The public methods mirror SincOversampler so the two are drop-in replacements. ModulatedOscillator now decimates through it: its SincRadius template parameter is gone (ModulatedOscillator<SampleType, OversampleFactor, CoeffType>), prepare() takes an optional HalfbandOversamplerDesign, and getLatencyInSamples() became an instance method that reports the design's latency after prepare()

  • SincOversampler (the single-stage polyphase sinc formerly named Oversampler) is several times faster and rejects images and aliases far better. Each channel now keeps its history contiguously in front of the staging buffer so every output sample is one dotProduct (SIMD for float/float and, newly, double/double via FloatVectorOperations), replacing the per-tap circular-buffer modulo and strided table reads; the kernels are prebuilt phase-major with the gain baked in. Both kernels use Kaiser β = 9 (was 5), and SincTable::applyKaiserWindow now spans exactly the kernel radius (it previously windowed over (SincRadius + 1) · OversampleFactor entries while the kernel was used out to SincRadius · OversampleFactor, cutting it off at about 8 % of the window peak and capping the rejection of SincOversampler and Resampler alike). The Oversampler2xFloat … Oversampler8xDouble aliases moved from radius 8 to radius 16, so their reported latency grows from 16 to 32 input samples, and Oversampler16xFloat / Oversampler32xFloat plus their Double variants were added

  • Added PrismSpectrum (yup_dsp/oscillators/), promoting the graphics example's Prism recipe into the library. It shapes a FourierSeries through a raised-cosine ridge comb in log2-harmonic space, magnitude companding, exact spectral pulse-width modulation and a quadratic phase dispersion, renormalizing to the source's sum |c|. A spectral tilt, an odd/even balance, a movable formant resonance and a pseudo-random phase scatter sit alongside those, each continuous through its own neutral value so it can be modulated without a step; scatter shares dispersion's rotation and so costs nothing extra. Every stage is a per-harmonic scale or rotation, so none can add a frequency and the backend's own bandlimiting is preserved. There is no transform behind it, so a shape can be re-derived once per audio block and the controls modulated; prepare() precomputes the per-harmonic log2 tables. The pulse-width depth fades in over the first squeezeFadeWidth, because the raw factor tends to a differentiator rather than to unity as the width falls, which would otherwise make zero a discontinuity

  • Added VAStateVariableFilter (yup_dsp/filters/), a topology preserving transform state variable filter modeled on Zavalishin's papers. One pass yields lowpass, highpass, constant-skirt and unity-gain bandpass, band shelf, notch, allpass and lowpass-minus-highpass outputs; resonance can be set in 0..1 or as Q, the shelf gain in dB, and the complex response is analytic

  • Added LFO (yup_dsp/oscillators/), an allocation-free control-rate oscillator with sine, triangle, sawtooth, square and seeded sample-and-hold shapes, a phase offset, and skip() for block-rate consumers

  • Added WaveformBank::refreshFrames(), replacing a prepared bank's coefficients without allocating and keeping the per-level harmonic limits prepare() chose, plus ModulatedOscillator::setBank() for adopting a bank rebuilt off-thread

  • ModulatedOscillator::Parameters moved to namespace scope as ModulatedOscillatorParameters (still aliased as the nested Parameters), and the modulation maths moved into detail::ModulatedOscillatorVoice so oscillators can compose over it without duplicating the phase map, sync and BLEP/BLAMP corrections

  • Added a waveform selector (sine, triangle, saw, square) and 16× / 32× sweep oversampling modes to the spectrum analyzer example; the non-sine shapes are naive, so the sweep oversampling modes show their aliasing suppression. The oversampled sweeps now generate at a multiple of the device rate and decimate straight to it through SincOversampler::beginGeneration() instead of passing through the radius-8 resampler, whose 54 dB stopband was setting the alias floor regardless of the oversampling factor

  • Reduced the graphics synthesizer example to a single Prism oscillator per slot, deriving each slot's spectrum once per block rather than once per voice so the shape controls can be modulated, and exposing squeeze, squash, tilt, odd/even, formant and scatter alongside ridges, color and dispersion. The sync modes moved onto Prism: the shaped series goes through SyncSpectralResampler at the slot, driven by the sync mode chooser and a new sync ratio knob, so unison satellites and the waveform preview follow it for free; the Modulated algorithm, its FM knobs and the algorithm chooser are gone from the example. The example then grew a per-voice VAStateVariableFilter stage with drive and keytracking, a second envelope, two per-voice LFOs (retriggered per note or free running) with displays following the newest voice, and an eight-slot modulation matrix on a second page: every source is per voice, and a voice derives a private spectrum only while a route reaches one of its spectrum parameters. Knob captions show the value while dragging, and Randomize now covers every voice property (oscillators, filter, envelopes, LFOs and a few live routes) while leaving volume, voice mode, glide and MIDI input alone. The example is split into examples/graphics/source/examples/audio/ (settings, engine, panels, modulation page)

  • Improved the graphics synthesizer example with a spectral Prism oscillator, octave and cents controls, poly/mono/legato modes, portamento, rendered voice-stealing tails, a revised instrument layout, and audio-load/overrun metering. Reduced callback and scope overhead, removed output waveshaping, and added regression coverage.

  • Added shared WaveformBank and oversampled ModulatedOscillator with through-zero FM, PM, phase distortion and fractional hard sync. Added direct generation to SincOversampler; fused spectral SIMD accumulation, amortized additive phasor trigonometry, and corrected sync bandwidth refresh and Nyquist boundaries.

  • Fixed the pulsar spectral resampler's alternating coefficient sign and corrected oscillator regression tests for fixed-size copies, startup crossfades, and spectral window leakage.

  • Added a TuningMap class (midi/yup_TuningMap.h): maps MIDI note numbers to frequencies under an arbitrary scale and key map, loading Scala .scl scale files and .kbm key map files via loadScale() / loadKeyMap() (which return a yup::Result and keep the previous tuning when a file fails to parse). isNoteActive() reports the notes a key map asks to retune, taken from the range in its header unless the file carries < first last lines, which declare it instead

  • FFTProcessor is now templated on the sample type - FFTProcessor<float> (the default) or FFTProcessor<double> - and every backend (PFFFT, Apple vDSP, Intel IPP, FFTW3 and the Ooura fallback, which now ships both a float and a double implementation) gained a native double-precision path. References to the nested scaling enum need qualifying, e.g. FFTProcessor<float>::FFTScaling::asymmetric

  • Fixed the PFFFT-backed FFTProcessor requiring its input and output buffers to be SIMD aligned: the transforms are now staged through buffers owned by the PFFFT backend and allocated with PFFFT's own aligned allocator, so the public API accepts buffers with any alignment (the double-precision real transform previously hit PFFFT's VALIGNED assertion when handed a plain std::vector<double>)

  • Fixed the double-precision Ooura FFT translation unit only building on GCC/Clang: it declared every internal helper (makewt, cftfsub, bitrv2, ...) inside the body of the functions that call them, and a block-scope declaration inside namespace yup declares a global function, so yup::cdft referenced a ::makewt that no one defined and the Windows link failed with 30 unresolved externals. The declarations now sit at namespace scope, matching yup_OouraFFT8g_float.cpp

  • Added the yup_dsp oscillator classes (oscillators/yup_FourierSeries.h, oscillators/yup_SyncSpectralResampler.h, oscillators/yup_AdditiveOscillator.h, oscillators/yup_WavetableOscillator.h, oscillators/yup_SyncOscillator.h), an alias-free synchronizing oscillator following Roth, Keller, Castaneda and Studer, "Alias-Free Oscillator Synchronization via Additive Synthesis" (DAFx26, paper 49): FourierSeries holds the coefficients and the waveform presets, SyncSpectralResampler rewrites them from a follower period ratio into hard, mirrored or pulsar sync, AdditiveOscillator and WavetableOscillator synthesize the result using only the harmonics below Nyquist, and the SyncOscillator facade offers both backends and refreshes them from a once-per-block update(). utilities/yup_DspMath.h also gained fillHarmonicPhasors

Graphics

  • Added Vector3, Matrix4 (column-major, with perspective / orthographic / look-at factories and rive::Mat4 converters) and Ray (viewport unprojection, plane and triangle intersection), plus MeshSurfaceMapper in the new meshes/ folder, mapping between viewport points and texture coordinates of any triangle mesh.

  • SVG drawables now prepare text layout, fitted image bounds, gradients, local dashed paths, marker placements and path bounds while parsing, avoiding repeated CPU work and failed image-resolution retries during rendering.

  • Lottie loading now prewarms static shape, mask, gradient and stroke-dash caches plus spatial motion-path lengths, so their first rendered frame does not perform one-time geometry or paint preparation; static gradients and masks are reused on subsequent frames.

  • Image::getWidth() and Image::getHeight() now return 0 on an invalid image instead of asserting and dereferencing null. Other accessors and pixel access still assert, as documented

  • GpuTexture now caches the backend texture views it hands out, keyed by view descriptor. A render pass previously allocated a fresh ore::TextureView for every attachment on every pass and for every sampled texture on every draw - all identical frame after frame - which on Metal made building the attachment descriptors cost more than creating the command encoder they were for. The cache needs no invalidation because a GpuTexture wraps one underlying texture for its whole lifetime: GpuCanvas and GpuTarget build a new GpuTexture whenever their backing changes

  • GpuRenderPass now encodes every draw of a pass into a single backend render pass instead of opening and closing a fresh one per draw. Each draw() / drawIndexed() previously created its own command encoder - two per frame in a simple scene, one per pipeline in a multi-pass effect - which on Metal made renderCommandEncoderWithDescriptor: one of the most expensive calls in a frame. The pass is opened lazily by the first draw and registered with the context, so beginning another one still auto-closes it, and GpuFrame::submit() closes a pass the caller left open before committing. Attachments must now be bound before the first draw, which is asserted

  • GpuFrame no longer blocks on the GPU when it goes out of scope. Ending a frame cost a full pipeline stall on every rendered frame, justified by the encoded passes holding raw pointers to the frame's transient resources - but the backends YUP ships construct their ore context with no GPUResourceManager, so post-submit lifetime is already handled by Metal's command-buffer retention, OpenGL's deferred name deletion, and D3D11/WebGPU reference counting. A frame now hands its texture views, samplers and pooled uniform buffers to the device tagged with a generation, and a generation is released once enough later frames have begun. GpuFrame::waitForGPU() is unchanged and remains the explicit opt-in before a CPU readback

  • GpuFrame now reports real currentFrameNumber / safeFrameNumber values to the ore context, which previously received a value-initialised descriptor leaving both at zero on every frame - so a manager-backed backend (ore's Vulkan and D3D12 paths) could never reclaim anything

  • Fixed the Metal GraphicsContext creating its own MTLCommandQueue instead of sharing the GpuDevice's, even though the device already exposed getDevice()/getCommandQueue() for that purpose. Metal orders work within a queue but never between two of them, so present work encoded on the private queue could run before the offscreen RHI work it samples had finished. Every other backend already shares a single command stream (D3D11 shares the immediate context, WebGPU shares device.GetQueue(), OpenGL has one current context)

  • Added GpuTexture::create() and GpuTexture::upload(), so textures can finally be allocated and filled directly instead of only being obtained from a GpuTarget / GpuCanvas. Covers every shape the backend layer supports - 2D, cube, 3D and 2D-array - with mip levels, MSAA sample counts and the full 24-format list (8-bit r/rg/rgba/bgra, 16- and 32-bit float, rgb10a2unorm, r11g11b10float, the depth/stencil formats and the BC/ETC2/ASTC block formats)

  • Added GpuSampler and GpuRenderPass::setSampler(): filtering, wrap mode per axis, mip filter, LOD clamp, anisotropy and comparison samplers. Slots the caller does not set keep the previous hardcoded linear / clamp-to-edge sampler, so existing pipelines are unchanged. A comparison-sampler binding now also gets a comparison function, without which the binding was invalid

  • Added GpuTarget::create (device, GpuTextureDesc) and GpuTarget::createFromTexture (device, texture, view), which give render-to-mip and render-to-cube-face. The create (device, width, height) overload still goes through the Rive 2D canvas allocator and is unchanged. Note that a mip- or layer-narrowed view is honoured for attachments on every backend but not for sampling on OpenGL / OpenGL ES, which falls back to the base texture - read a specific mip with an explicit textureLod() and a LOD-clamped GpuSampler instead

  • Added GpuDevice::isFormatSupported(), isFormatRenderable(), isAnisotropicFilteringAvailable() and getMaximumSampleCount(). Float colour attachments are extension-gated on OpenGL ES / WebGL2, so a caller can now degrade instead of rendering black

  • GpuRenderPass now binds the attachments GpuPipelineOptions was always able to describe: up to four colour attachments via setColorAttachment() (colorTargets[4] previously documented MRT but the pass hardcoded a single attachment), a depth/stencil attachment via setDepthStencilAttachment() (depthStencil.enabled previously compiled a pipeline that no pass could satisfy), and an MSAA resolve destination via setResolveTarget()

  • GpuRenderOptions gained explicit GpuLoadOp / GpuStoreOp per attachment, plus GpuDepthStencilOptions for depth and stencil. The { bool clear, GpuColor } form is kept as a shorthand. Attachments are now cleared before the first draw of a pass only, so several draws can accumulate into one surface and share one depth buffer instead of each draw wiping its predecessors

  • GpuRenderPass bind groups now only carry the slots the bound pipeline's layout actually declares. Binding state outlives a single draw, so a slot bound for an earlier draw with a different pipeline used to be handed to the backend anyway, which rejects the whole group - silently dropping every binding in it. A bind group that still fails to build is now reported instead of leaving the draw reading whatever was bound before it

  • GpuRenderPass::draw() and drawIndexed() now forward instanceCount, firstVertex, firstInstance, firstIndex and baseVertex, so instanced drawing works - GpuVertexStepMode::instance was previously declarable but could never advance past instance 0. Added setViewport(), setScissorRect(), setStencilReference() and setBlendColor(), all sticky across the draws of a pass

  • GpuPipelineOptions gained GpuColorTarget::writeMask (GpuColorWriteMask), the remaining ore blend factors (srcAlphaSaturated, blendColor, oneMinusBlendColor) and the remaining vertex formats (sint8x4, uint16x2, sint16x2, unorm16x2, snorm16x2, uint16x4, sint16x4, float16x2, float16x4, uint32)

  • Fixed shader binding maps hardcoding every texture binding as a non-multisampled 2D float texture: the reflected image dimension, arrayed flag, depth flag and multisampled flag are now carried through, so a textureCube, texture3D, texture2DArray or depth texture no longer fails bind-group-layout validation on WebGPU. backendSpace is also now set to the binding's group, which was wrong for any set != 0

  • Fixed GpuSamplerDesc::label dangling: GpuSampler::create() copied the raw const char* into a retained member and handed it back through the public getDescription(), so any caller passing a temporary got a dangling pointer for the sampler's lifetime. Both it and GpuTextureDesc::label are now String

  • Fixed GpuPipelineCache not hashing GpuColorTarget::writeMask, so two pipelines differing only in write mask collided on the same cache key even though the mask is plumbed into the pipeline

  • Fixed the pipeline entry-point names being borrowed rather than owned: ore::Pipeline keeps a shallow copy of the PipelineDesc and dereferences its entry-point strings at draw time, while compileFromBundle() pointed them at a local String. The compiled GpuPipeline now owns them, alongside the vertex layouts it already copied

  • Fixed the WebGPU compute path ignoring the shader source length and requiring NUL-terminated code, which is not what an RSTB blob is

  • A render pass that binds a pipeline incompatible with its attachments (colour format, sample count or depth presence) now asserts and logs the backend's own diagnostic, instead of silently rendering nothing

  • Added a PBR IBL example to the graphics demo: a procedural sky baked into a float cube map face by face, convolved into an irradiance cube, prefiltered into a roughness mip chain, plus a split-sum BRDF lookup table and a CPU-uploaded albedo / normal map, shading an instanced grid of spheres against a real depth buffer

  • Fixed GpuDevice::isComputeAvailable() reporting from a runtime GL version probe on WASM / WebGL, where no GL compute implementation is compiled in at all; its documentation also claimed compute was available on D3D12 and Vulkan (neither backend exists) and unavailable on OpenGL (where it is implemented)

  • Fixed native/yup_GpuDevice_dawn.cpp guarding its whole body on RIVE_DAWN while the module includes it under YUP_RIVE_USE_DAWN, so the Dawn device could compile to nothing

  • Added GpuFrameDescriptor (rhi/yup_GpuTypes.h), a field-for-field mirror of rive::gpu::RenderContext::FrameDescriptor that keeps GpuDevice::beginOffscreen's public signature free of a Rive type. GpuCanvas::beginDraw() gained an optional const GpuFrameDescriptor& parameter (defaulting to today's behaviour: clear to transparent black, no msaa), giving callers control over msaaSampleCount, ditherMode (new GpuDitherMode enum), loadOp and clearColor for the offscreen 2D frame it opens

  • GpuCanvas::beginDraw() and the offscreen Graphics constructor take a scale (canvas pixels per logical unit): getContextScale() reports it and the default drawing area is the canvas size divided by it.

  • Fixed GPU compute silently stalling after a few frames on OpenGL with some drivers (AMD desktop GL): compute now runs on a dedicated, unshared GL context (GpuDevice::Options::computeContextActivator, routed through GpuDevice::runOnComputeContext()) which exclusively owns every compute resource — pipeline compilation, dispatches, storage buffer create/update/readback and deletion — falling back to the rendering context when unavailable. The GL compute pass also saves and restores the program and GL_UNIFORM_BUFFER bindings it touches, so it can no longer desync Rive's cached GL state

  • Fixed GL storage buffers being deleted right after creation (moving a GpuBuffer::Impl copied the plain GL buffer name, so the moved-from object's destructor freed the just-created buffer) and a crash when releasing GPU buffers or compute pipelines after their window closed (GL releases are routed through the owning device and skipped once the window's contexts are gone)

  • Fixed Emscripten randomly rendering nothing or freezing the tab: the render-thread rework unbound the GL context between frames so the render thread could take it, but on Emscripten rendering is timer-driven on the single browser thread and message-thread work between frames (image decodes, font atlas uploads) still issues GL calls — with no context current those throw in JS and kill the main loop. With timer-driven rendering the window context now stays permanently current

  • Fixed Emscripten freezing the browser tab when a demo requested a headless GL compute device: with no current WebGL context every GL call throws in JS, killing the requestAnimationFrame main loop. The GL GpuDevice now fails construction gracefully when no WebGL context is current and isComputeAvailable() reports false on a device whose GL initialization failed. GpuAudioProcessingDemo also sized its CPU ring buffers only when GPU compute was available, so the no-GPU audio path wrote and read out of bounds

  • Fixed SDLComponentNative::renderFrame() on Emscripten encoding a frame with a zero-sized render target, which trips Rive's beginFrame() assertion — fatal there, since an assertion abort throws inside the requestAnimationFrame callback and permanently kills the browser's main loop. It now skips rendering entirely until a non-zero content size has been observed

  • Added a native WebGPU GraphicsContext backend for Emscripten via the Emdawnwebgpu port (RIVE_WEBGPU=2 + --use-port=emdawnwebgpu, enabled with the ENABLE_EMSCRIPTEN_WEBGPU parameter of yup_standalone_app), rendering Rive content through the browser's WebGPU API without Dawn

  • Fixed GpuFrame::begin() aborting on the Emscripten WebGPU backend: the WGPU context now creates and submits its own command encoder when no external one is provided, matching the Metal/GL/D3D11 self-managed frame model

  • Fixed a crash on Windows when creating any native window: the D3D11 GpuDevice was built with an already moved-from ID3D11Device, and the Direct3D GraphicsContext created a second device whose swapchain textures could not be used by the render context. Both now share a single ID3D11Device

  • Fixed the Emscripten WebGPU GraphicsContext never storing its surface size, leaving the offscreen copy at 0x0

  • Fixed GpuDevice::updateBuffer() failing for every vertex, index and uniform buffer on the WebGPU, Dawn and D3D11 backends: those overrides handled native storage buffers only and returned false instead of delegating ore-backed buffers to the base class, the way the Metal and OpenGL overrides do

  • Implemented GpuDevice::readBuffer() for D3D11, which previously reported isComputeAvailable() but had no override, so every storage buffer readback silently failed through the base class. It copies into a cached D3D11_USAGE_STAGING buffer on the immediate context (ordered after the dispatch) and maps it for reading

  • GpuComputePass on D3D11 now unbinds the UAV slots it bound when the pass finishes, so a storage buffer is no longer left bound for writing while a later readback or draw reads it

  • Fixed GpuDevice::readBuffer() never succeeding on the Emscripten WebGPU backend: it mapped its staging buffer with WGPUCallbackMode_AllowProcessEvents and then tested the result in the same call, but WebGPU buffer mapping only resolves through the JavaScript event loop, so the callback could not have run. The WGPU backend now pipelines the readback over a ring of three staging buffers using WGPUCallbackMode_AllowSpontaneous, which completes on its own between main-loop ticks - no ASYNCIFY needed

  • GpuDevice::readBuffer() is no longer documented as unconditionally blocking. Whether it blocks is a property of the backend: Metal, D3D11 and OpenGL read back in lockstep and fill the destination every call, while WebGPU cannot map synchronously and so trails the GPU by a frame or two. Callers must now own the destination across calls and treat a false return as "no new data yet" rather than an error - the previous contents stay valid

  • ComputeParticlesDemo: keeps drawing the last particle snapshot on frames where no new one has landed, so it renders on the Emscripten WebGPU backend instead of showing nothing. The status label reports the landed-snapshot count alongside the frame count

  • Component's effect path now reuses its offscreen GpuCanvas across frames while the component size is unchanged, instead of allocating (and freeing) a full-size render target every frame. On a size change the outgoing canvas is released before the replacement is created, so its RenderContext lease returns to the pool rather than forcing a second context to be reserved permanently

  • ComponentEffectsDemo: shader effects now share a common base that compiles the pipeline at most once instead of retrying a failed compile on every frame, reports the compile error in the status label and on the console, and shows the CPU time spent applying the effect next to the paint time

  • New Layout example (examples/graphics): exercises the FlexBox and Grid layout containers with pages for direction, wrap, justify-content, align-items, align-content, flex grow/shrink/basis, align-self, order, margins/gaps, percentage sizing, min/max constraints, grid track sizing (px/fr/auto), explicit placement and spans, auto placement, grid alignment, and nested flex/grid compositions

  • Added a Toast Notifications demo to the graphics example app demonstrating the ToastNotification utility: a simple sendNotification, a rich ToastTemplate (attribution, actions, scenario, duration, expiration, event callbacks), an image template, and hide/clear

  • Added a Fluid Simulation demo to the graphics example app: an incompressible Navier-Stokes solver (splat, vorticity confinement, divergence, pressure projection, gradient subtraction, semi-Lagrangian advection) ported from Pavel Dobryakov's WebGL Fluid Simulation, running every pass as a fullscreen-triangle GpuPipeline fragment shader over rgba8 ping-pong GpuTarget surfaces (sim fields packed as 16/24-bit fixed point) with a simplified bloom (original threshold/soft-knee prefilter into a small surface, blurred and upsampled with linear filtering, ordered dithering to hide 8-bit banding on fades) and shading composite presented via Graphics::drawTexture — no CPU readback. Pipelines compile incrementally with a progress bar, and a calibration pass keeps sample orientation consistent across backends

  • SDLComponentNative now renders each window on its own dedicated render thread instead of the message thread: the component-tree walk runs under a MessageManagerLock while GL command submission and buffer swap happen unlocked, so multiple windows no longer serialize their frame rendering (and vsync waits) on the message thread

  • Fixed SDLComponentNative::repaint() calling -[NSWindow screen] (via getSize() → getWindowUnitsPerPoint() → SDL_GetDisplayForWindow()) from the render thread, which macOS's Main Thread Checker flags since AppKit requires that call on the main thread. It now uses the already up-to-date screenBounds cached by the main-thread window event handlers instead of querying the display live

  • Fixed SDLComponentNative::runWithGraphicsContext() never invoking its callback on non-OpenGL desktop backends (Metal, Direct3D), silently dropping the work. Component::renderSubtreeOffscreen() routes through this hook, so any component with a ComponentEffect set (e.g. ComponentEffectsDemo) rendered nothing on Metal/D3D; SDLComponentNative::runWithComputeContext() falls back to the same hook when no dedicated compute context exists, so GPU compute work initiated off the render thread was silently dropped there too

  • Fixed Android crashing on any touch input after the render-thread rework: SDL invokes event-watch callbacks on the thread that pushes the event — the OS UI thread on Android, not the message thread — so the whole input/window-event pipeline mutated the component tree concurrently with the message and render threads. Events arriving off the message thread (window, input, drop and display events) are now marshaled onto it via MessageManager::callAsync, with SDL-owned string payloads (SDL_EVENT_TEXT_INPUT, SDL_EVENT_DROP_FILE/DROP_TEXT) deep-copied since SDL frees them when the watch returns

  • Mobile app lifecycle is now handled explicitly: SDL_EVENT_WILL_ENTER_BACKGROUND stops the render thread synchronously inside the event watch (the GL surface can be destroyed as soon as the callback returns, and SDL may block the message thread until the app resumes), and SDL_EVENT_DID_ENTER_FOREGROUND restarts rendering from the message thread. ~SDLComponentNative also removes its event watch before any teardown, so no other thread can dispatch events into a half-destroyed window

  • Offscreen GPU work initiated from the message thread (component snapshots/effects via runWithGraphicsContext()) now serializes against the render thread's frame submission on every backend: renderFrame() submits GraphicsContext::end()/swap outside the MessageManagerLock, so Metal/D3D needed the same context lock the OpenGL path already used

  • Fixed GPU resources leaking when released off the render thread: dropping a cached component canvas, snapshot, or GPU-backed Image on the message thread issued GL deletes with no context current (the render thread owns it now). GpuTarget and GpuTexture teardown routes through the new GpuDevice::runOnGraphicsContext(), which binds the rendering context on GL (mirroring the existing runOnComputeContext() buffer-release path); the window destructor also keeps the GL context current while destroying its renderer and graphics context

  • Fixed the window's GPU context activators dangling when GPU buffers or textures outlive their window: the activators now share a refcounted guard with SDLComponentNative (the same lock that serializes GL context access), so a release after the window is gone locks, sees a null window and bails out instead of calling into freed memory

  • Fixed the render thread overflowing its stack on complex paths: user paint() code now runs on a secondary thread whose default stack is far smaller than the main thread's (512kB vs 8MB on macOS), and Rive's GrTriangulator recurses hundreds of frames deep on many-point filled paths (e.g. SpectrumAnalyzerComponent's spectrum). The render thread is now created with an 8MB stack to match the main thread paint() previously ran on

  • Fixed the render loop waiting forever when framerateRedraw is above 250: the frame budget minus the 4ms pacing margin went negative, which WaitableEvent::wait() treats as wait-without-timeout

  • ComponentNative::setDesiredFrameRate() can now change a window's target framerate while it is rendering, and the new Options::withUnfocusedFramerateRedraw() throttles a window to a lower rate whenever it does not have keyboard focus, for example while it sits behind a modal window. The unfocused rate only ever lowers the rate, and withUpdateOnlyFocused() still takes precedence by stopping rendering entirely. getDesiredFrameRate() keeps reporting the rate that was asked for rather than the throttled one

  • WebAssembly frames are now produced directly from requestAnimationFrame rather than from a Timer, removing a cross-thread handshake from the frame path: emscripten builds always pass -pthread, so the timer countdown is kept on the timer worker thread and reaches the main thread as a posted message that is drained on the next animation frame. The animation frame is also the only moment the browser presents, so a target rate below the display rate is now met by rendering every Nth callback instead of against a deadline, which can land just after a refresh boundary and cost a whole interval. SDLComponentNative::renderDrivenByTimer is removed in favour of YUP_EMSCRIPTEN ifdefs.

  • Added Timer::getTimerFrequencyHz(), which reports the rate from the sub-millisecond interval instead of round-tripping it through whole milliseconds the way 1000 / getTimerInterval() does. SpectrumAnalyzerComponent::getUpdateRate() now uses it and returns exactly the rate that was set

  • Timer now keeps its period and countdown at sub-millisecond resolution and carries the overshoot of a late callback into the next period, so a rate that is not a whole number of milliseconds is honoured on average: startTimerHz (60) previously truncated 1000 / 60 to a 16ms interval and ran at 62.5Hz. The timer thread also measures elapsed time with Time::getMillisecondCounterHiRes() rather than the whole-millisecond counter. getTimerInterval() keeps its millisecond contract and now rounds, so a 60Hz timer reports 17 rather than 16, and a timer faster than 1kHz reports 1 rather than 0

  • Fixed timers never firing in a WebAssembly build without pthreads: the countdown is done by the timer thread, and the fallback that counted down on the message loop instead was behind #if YUP_EMSCRIPTEN && ! defined(__EMSCRIPTEN_PTHREADS__), which never held for a target built with yup_add_standalone_app, the only place that flag is added. It is now selected at runtime by whether the timer thread is actually running, which also covers any other platform where it cannot be started

  • The SDL render thread now paces frames on an absolute schedule instead of sleeping for whatever is left after each frame, so a frame's own overshoot is no longer carried into the next one, and it waits for a repaint request only until the point where renderFrame() still fits inside the frame, using a rolling average of the measured render cost rather than a fixed 4ms margin (at 60Hz the old margin left 12.7ms of waiting, which plus any real render work overran the 16.7ms frame). On Apple the render thread is also started as a realtime thread, since outside the Mach time-constraint class the kernel coalesces timer expiries into a window of roughly 25% of the requested sleep

  • Fixed glslang process initialization racing between concurrent ShaderTranspiler constructions (per-window render threads can now compile shader effects while a compute pipeline compiles elsewhere): the refcount was atomic but did not make a second thread wait for InitializeProcess() to complete; init/finalize now run under a lock. Also made the offscreen context pool's leased flag atomic in the GL and Metal devices, since RenderableTarget destructors can run outside the context lock that serializes acquisition

  • Fixed WaitableTimer's non-Windows fallback overshooting frame deadlines by several milliseconds on macOS: it blocked for the entire remaining wait on a single condition_variable::wait_until, whose wake time is subject to OS scheduling / timer-coalescing latency. It now blocks for the bulk of the wait and closes the last few milliseconds with a tiered busy-wait against Time::getMillisecondCounterHiRes(), restoring the precision the pre-WaitableTimer implementation had

  • SDLComponentNative now caches its window units-per-point value: getWindowUnitsPerPoint() is only called when the window is created and on resize/dpi/screen-change events (SDL_EVENT_WINDOW_DISPLAY_CHANGED, DISPLAY_SCALE_CHANGED, PIXEL_SIZE_CHANGED, MOVED), and every other read (including from the render thread) uses the cached value, so no SDL window method is ever invoked from the render thread

  • New ComponentNative::vsync flag / Options::withVSync() (off by default): synchronizes presentation to the display refresh (Metal displaySyncEnabled, GL swap interval, D3D Present (1)). With vsync the present paces the render thread: after a frame that was presented and waited for the refresh, the next one starts without the software deadline, while frames that paint nothing (or backends whose present doesn't wait, like Dawn) keep it so the loop can't spin

Rive Runtime Bump

  • Rive runtime bumped from v0.1.62 to v0.1.155

RHI (#129 and #130)

  • GraphicsContext GPU context integration. New GraphicsContext::isGpuAvailable() capability probe; gpuContext() is retained but documented @internal as the single backend bridge.
  • New GpuTexture class (rhi/yup_GpuTexture.h): opaque reference-counted GPU texture wrapping rive::gpu::Texture or rive::gpu::RenderCanvas. Obtained from GpuCanvas::asTexture() or constructed internally by Image::fromTexture().
  • New GpuTarget class (rhi/yup_GpuTarget.h): low-level render-pass-only offscreen GPU surface (create, beginRenderPass, asTexture, asImage, readPixels). Its backing texture is allocated from the context's main render context, so it does not reserve a dedicated rive::gpu::RenderContext - use it for custom GpuPipeline work (e.g. post-process passes) that needs no 2D drawing.
  • New GpuCanvas class (rhi/yup_GpuCanvas.h): consolidated backend-agnostic offscreen GPU surface that now composes a GpuTarget (over a RenderableTarget) and creates a non-owning Graphics lazily only when 2D drawing is requested.
  • Python bindings now expose GpuColor as yup.GpuColor (backing GpuRenderOptions.clearColor), comparable with yup.Color.
  • Python bindings now expose GpuLoadOp / GpuStoreOp and GpuRenderOptions.loadOp / .storeOp. GpuRenderOptions.clear is kept as a bool view of loadOp, so GpuRenderOptions(True, color) and opts.clear still read the same as before.

RHI module extraction & GpuDevice

  • New yup_rhi module: the GPU abstraction layer extracted from yup_graphics into its own module (depends on yup_core, yup_shading, rive_renderer). All RHI classes (GpuFrame, GpuPipeline, GpuBuffer, GpuTexture, GpuTarget, GpuRenderPass, GpuPipelineCache) now live in yup_rhi. yup_graphics depends on yup_rhi for GPU access.
  • GpuDevice: new reference-counted GPU device abstraction (was GpuContext). Owns the native GPU device and command queue without requiring a window - can be used for headless GPU compute (e.g. audio DSP on the GPU). Created via GpuDevice::create(GpuPlatform, Options). All RHI factory methods (GpuFrame::begin, GpuPipeline::compile*, GpuBuffer::create, GpuTarget::create) now take GpuDevice::Ptr for safe shared ownership.
  • GpuPlatform enum: standalone platform enum (Headless, Metal, Direct3D, OpenGL, OpenGLES, WebGPU) replacing the nested GpuDevice::Api. GraphicsContext::getPlatform() returns it directly - no typedef alias.
  • GpuColor struct (rhi/yup_GpuTypes.h): lightweight 4-component GPU color for render options. Implicitly constructable from any type with getRedFloat()/getGreenFloat()/getBlueFloat()/getAlphaFloat() (e.g. yup::Color), so GpuRenderOptions { true, Colors::transparentBlack } works without code changes.
  • GraphicsContext simplified: wraps a GpuDevice::Ptr (obtained via getGpuDevice() returning GpuDevice::Ptr). Offscreen target management (createOffscreenTarget, beginOffscreen, endOffscreen, readOffscreenPixels) delegated to GpuDevice. Factory accepts optional GpuDevice::Ptr to share an existing GPU device.
  • Backends: GpuDevice has native implementations for all platforms (Metal, OpenGL, Direct3D 11, Dawn, WebGPU/Emscripten, Headless). OpenGL backend probes GL_VERSION at runtime to detect compute shader support (GL ≥4.3 / GLES ≥3.1).
  • ::Ptr safety: all RHI types that own resources (GpuPipeline, GpuBuffer, GpuTexture, GpuTarget, GpuCanvas) are reference-counted with ::Ptr. Factory methods take GpuDevice::Ptr to keep the device alive for the resource's lifetime. GpuFrame is move-only stack RAII and takes GpuDevice& (no ownership).

Compute Shaders & GPU Audio

  • New GpuComputePipeline class (rhi/yup_GpuComputePipeline.h): an immutable compiled compute pipeline that bypasses ore to go directly to the backend-native API (Metal MTLComputePipelineState, D3D11 ID3D11ComputeShader, WebGPU/Dawn wgpu::ComputePipeline, OpenGL GL_COMPUTE_SHADER). compile(ctx, source, GpuWorkgroupSize), compileFromBundle(ctx, ShaderBundle), and compileFromGlsl(ctx, glsl) (when YUP_ENABLE_SHADER_TRANSPILER = 1) all return ResultValue<GpuComputePipeline::Ptr>.
  • New GpuComputePass class (rhi/yup_GpuComputePass.h): move-only RAII compute dispatch encoder (GpuComputePass::begin(device)). Binds a GpuComputePipeline, storage buffers (setStorageBuffer), uniform buffers (setUniformBuffer), and textures (setTexture), then dispatches workgroups via dispatch(gx, gy, gz).
  • GpuBuffer extended with GpuBufferType::storage: native storage buffer creation for each backend (Metal MTLBuffer, D3D11 structured buffer + UAV, WebGPU Storage buffer, OpenGL GL_SHADER_STORAGE_BUFFER). Storage buffers are bound to GpuComputePass::setStorageBuffer().
  • GpuDevice backends expose native compute handles: getDevice()/getCommandQueue() (Metal), getD3DDevice()/getD3DDeviceContext() (D3D11), getWgpuDevice()/getWgpuQueue() (WebGPU/Emscripten), getBackendDevice()/getDevice()/getQueue() (Dawn).
  • GpuAudioProcessingDemo example: real-time GPU-accelerated audio effect (gain + soft clipper) using compute shaders. Captures live audio via AudioIODeviceCallback, uploads to GPU storage buffers, dispatches a compute shader, and reads back processed audio - all on the audio I/O thread.
  • New GpuDevice::updateBuffer(): writes new data into an existing storage buffer without reallocating it (Metal contents memcpy, D3D11 UpdateSubresource, WebGPU/Dawn WriteBuffer, GL glBufferSubData). Fixes GpuAudioProcessingDemo reallocating its input storage buffer every audio callback, which caused audible stutter. The gain/mix parameters remain a uniform buffer (as before) - that path is unaffected and its small per-dispatch allocation is negligible next to the audio-block-sized buffer this fix removes.
  • Fixed ShaderTranspiler's MSL backend assigning storage/uniform buffer indices via spirv-cross's own auto-incrementing scheme instead of the shader's declared layout(binding=N): added CompilerMSL::Options::enable_decoration_binding = true so the compiled [[buffer(N)]] index always matches the declared binding, matching what GpuComputePass's native dispatch (which binds slots as group*16+binding with no reflection indirection) requires. This was silently producing zero output from any Metal compute shader with more than one storage/uniform buffer, including GpuAudioProcessingDemo.
  • Metal GpuDevice/GpuComputePass calls now wrap their Objective-C work in @autoreleasepool blocks - without one, real-time callers (e.g. an audio thread with no ambient pool) accumulated command buffers/encoders indefinitely.
  • Metal GpuDevice/GpuComputePass calls now wrap their Objective-C work in @autoreleasepool blocks — without one, real-time callers (e.g. an audio thread with no ambient pool) accumulated command buffers/encoders indefinitely.
  • Fixed ShaderTranspiler emitting ESSL 3.00 for every stage, which made spirv-cross reject compute shaders on OpenGL ES ("At least ESSL 3.10 required for compute shaders") — compute stages now target ESSL 3.10 (#version 310 es) while vertex/fragment stages keep ESSL 3.00.

Image Formats

  • Added TIFF read/write support (TiffImageFormat) via libtiff: RGB, RGBA, Grayscale at 8/16-bit; multi-page reading; DPI and EXIF/ICC/XMP metadata extraction.
  • Added TGA read/write support (TgaImageFormat): uncompressed and RLE-compressed truecolor and grayscale variants; RGB and RGBA output with alpha channel preservation.
  • Added animated WebP encoding and decoding support to WebPImageFormatWriter and WebPImageFormatReader: per-frame metadata (canvas dimensions, frame count, loop count, per-frame delays, dispose/blend modes) and frame decompression with manual compositing.
  • Added animated PNG (APNG) encoding and decoding support to PngImageFormatWriter and PngImageFormatReader: manual chunk-level parsing of acTL/fcTL/fdAT chunks for animation metadata, per-frame libpng decoding via synthetic minimal PNG construction, and canvas compositing supporting all three APNG disposal operations (none, background, previous) and both blend operations (source, over).
  • ImageFormat::Options struct controls metadata extraction: .withMetadata(true) enables text metadata and DPI; .withRawChunks(true) enables raw binary chunks (EXIF, ICC, XMP). When both are false (the default), ImageMetadata is not allocated - true zero overhead.
  • Introduced a ref-counted ImageMetadata object (ImageMetadata::Ptr) attached to Image and ImageFormatReader::metadata. DPI, text entries, and raw binary chunks are all accessed through the metadata object only when requested via Options.
  • Lossless roundtrip tests for all formats (BMP, PNG, WebP, TGA, TIFF, PPM, GIF) now verify pixel-perfect fidelity after write→read; animated roundtrip tests for GIF, WebP, and PNG verify per-frame pixel integrity.
  • StyledText::TextModifier::appendText() gained a Color overload that creates (and caches per color) a solid fill paint, and Graphics::fillFittedText() now honors per-run style paints when every run carries one - enabling syntax-colored text. Single-color StyledText usage is unchanged. Font gained isEmpty().
  • New Font static loaders: Font::loadFontFromData(), Font::loadFontFromFile(), Font::loadFontFromFirstAvailableFile(), Font::loadSerifSystemTextFont() and Font::loadMonospaceSystemTextFont(), all returning ResultValue<Font> (wasOk() / failed() / getValue()). The former theme-local system font lookup helpers moved into Font; macOS/iOS use the CoreText system UI fonts, other platforms try well-known system font files.
  • ImageFormatReader and ImageFormatWriter gained a deleteSourceWhenDestroyed constructor parameter, defaulting to true so existing callers are unaffected. It lets a format be handed a stream it did not open and either take ownership of it - deleting it on destruction, as the C++ path has always done - or decline and leave the caller to own it, in which case the reader's / writer's owning pointer is released rather than deleted. The reader overload that accepts the caller's InputStream alongside the flag is what lets ImageFormatManager::createReaderFor hand a Python format the stream it opened without leaking it

UI

  • Input follows what is displayed: ComponentEffect::displayToContent() / contentToDisplay() let distorting effects route pointer input, and the new Component::getChildPointFromLocal() / getLocalPointFromChild() hooks, used by every input path, let a parent present children through a custom projection. Component::setManuallyComposited() and renderToTexture() let a parent render a live child subtree itself, for example on a 3D surface. When painted content moves under a stationary pointer, the SDL window now re-evaluates the pointer after the frame and sends the enter/exit, move or drag events it implies. ComboBox popups open inside the closest transformed or manually composited ancestor (Component::getPopupParentComponent()), so they are presented with the widget.

  • Fixed a component's opacity being applied twice when it has an effect or is cached to texture: the offscreen content is rendered opaque and the opacity is applied once when compositing.

  • Fixed Component::getTransformToScreen() adding a desktop window's position twice and applying the root's own transform; it now matches localToScreen().

  • Component::findComponentAtForMouseEvent() maps the point into children through transforms, effects and getChildPointFromLocal(), like findComponentAt(), so mouse input reaches transformed, warped and manually composited components.

  • Effect and cached-to-texture canvases are rendered at the display scale, so they stay sharp on high-density displays. Component::renderToTexture() takes a scale (pass g.getContextScale()). Snapshots stay at one pixel per point.

  • Fixed animation repaint requests bypassing the render deadline, which produced uneven frame submission and could run the Metal render loop uncapped when VSync was enabled. Repaint events now wake the render thread without allowing a frame before its scheduled deadline, and missed deadlines skip directly to the next future frame

  • Animation precomposition canvases are now leased from a bounded per-frame scratch pool. Their cache previously used the sampled animation phase as a persistent key, allocating another full-size GPU canvas for each new animation frame and retaining all of them until the renderer was reset

  • Component gained the callbacks it was missing for reacting to its surroundings: childBoundsChanged (child) on a parent whose direct child moved or was resized (a single setBounds() reports once, not once for the move and once for the resize), parentSizeChanged() on each direct child after its parent's resized() has run, indexInParentChildrenChanged (oldIndex, newIndex) on a component whose z-order position changed, and focusOfChildComponentChanged (child, cause) on every ancestor of a component that gained or lost the keyboard focus

  • Component::keyStateChanged (key, isDown) reports a key going down or coming up even when the press never reaches keyDown(), and does not retrigger on auto-repeat, which makes it the place to track a held modifier or chord key. Component::modifierKeysChanged (modifiers) reports the modifier state changing, including when it changes during a mouse gesture rather than a key event

  • Component::hitTest (x, y) decides whether a local point counts as being inside a component, and overriding it overrides hit testing for the mouse: a non-rectangular widget can return false for its transparent corners and let events through to whatever is behind it. An override can only take area away, since a point outside a component's bounds is rejected by its parent before hitTest is consulted

  • Fixed mouse ups going missing. The windowing backend polled the live OS button state on a timer and dropped the whole gesture whenever it saw no button down, but that state changes the moment the button is physically released, while the matching SDL_EVENT_MOUSE_BUTTON_UP is still queued: a tick landing in between left the following mouse up with no component to deliver it to, so the component stayed pressed and never saw the click. The poll is gone; a gesture whose release was genuinely consumed elsewhere - a native drag session runs its own event loop and swallows it - is now ended explicitly through the new ComponentNative::cancelCurrentMouseGesture(), which delivers the missing mouse up rather than silently forgetting it.

  • Fixed a window losing mouse capture after a drag. setGlobalMouseCaptureActive() and the window's own captureMouse shared a single flag to track a reference on one process-wide capture count, so a drag starting from a window that already captured released that window's capture when it ended, and a setVisible() during a drag released the drag's. Each now owns its own reference.

  • Drag and drop can now leave the application. DragOptions::withExternalDragAllowed (true) opts a source in, and the manager hands the gesture to the platform once the pointer is over none of our windows: on macOS that runs a native AppKit drag session carrying files, text and PNG images, so a YUP drag can be dropped on another application. An image payload is already PNG encoded by DragAndDropData, so it goes onto the pasteboard as it is. Windows and X11 have no implementation yet and keep such a gesture in the app - and note that a source cannot be both in-app and external, because crossing between two windows is indistinguishable from leaving the application.

  • New "Drag and Drop" example in examples/graphics: two trays of tiles that drag between each other and into a second window opened from the demo (live tiles are registered process-wide, so a tile can move across windows), a drop zone that accepts files and text dragged in from another application and lists what arrived, and a list whose rows drag out with multiple selection - starting a drag on one of several selected rows carries the whole selection.

  • ListBox is now a drag source: dragging a row asks the model for getDragSourceDescription (selectedRows) and, when that returns anything other than a default-constructed var, begins a drag carrying it. A string description is mirrored into the text MIME type as well, so a target that reads only MIME data still sees something. New createDragSourceComponent (selectedRows) produces the drag image and can be overridden; the default is a circle carrying the number of dragged rows. setDragSourceEnabled (false) makes a list undraggable without consulting its model at all.

  • Fixed ListBox multiple-selection behaviour: a plain click now replaces the selection, having previously added to it because selectRow() only ever adds in multiple selection mode. Shift extends a range from the last click without shift and keeps that anchor across further shift-clicks, so repeated shift-clicks grow or shrink one range instead of creeping along a row at a time, and command/control toggles an item. The range branch also repaints now, which it never did - the rows really were selected, but nothing showed it.

  • Fixed modifiers being misread on every mouse and touch event: the windowing backend built a KeyModifiers straight from SDL's SDL_GetModState() bitmask rather than through the mapping the keyboard path uses. SDL's bit layout does not overlap KeyModifiers's masks, so control, command and alt never registered at all, and shift only appeared to work because SDL's left-shift happens to sit on the same bit as shiftMask.

  • Fixed translucent windows never being cleared, and the compositor discarding their alpha anyway. The renderer only used LoadAction::clear for a window that renders continuously, so a window that draws on demand kept whatever the drawable held before; it now clears whenever the clear colour is not opaque. Alongside that, the Metal layer no longer forces opaque, a Windows swap chain created for a composed window requests DXGI_ALPHA_MODE_PREMULTIPLIED, and desktop OpenGL asks for an alpha channel as GLES already did. A window created with ComponentNative::Options::withTransparent (true) now composites properly rather than showing its un-cleared corners.

  • Drag and drop is now opt-in mixins plus an app-global session rather than part of Component (see Breaking changes). DragAndDropSource starts a drag from a payload and an optional ghost, DragImageComponent is the borderless window that renders that ghost, and DragAndDropManager owns the drag in flight: it listens for global mouse events for its duration, finds the component under the cursor across every native window (Desktop::findComponentAt), and drives the enter/move/exit and drop dispatchers. The performed action defaults to copy, or move with Shift held. DragAndDropTargetComponent is a convenience for a component that is also a target. Drags that pass over a window from outside the application arrive through the same machinery; note that the OS reports no payload until the drop itself, so a target that wants to react while such a drag hovers must accept an empty payload.

  • New ComponentNative window options withTransparent(), withAlwaysOnTop() and withFocusable(), mapping onto SDL's transparent, always-on-top and not-focusable window flags, plus ComponentNative::getComponent() and setGlobalMouseCaptureActive() so a window can keep receiving mouse input while the pointer is outside it.

  • Fixed Component::addChildComponent() (and therefore addAndMakeVisible()) not detaching a component from its previous parent. Reparenting left the component in the old parent's child list as well, so it was still painted and laid out there while getParentComponent() pointed at the new one, and a further reparent added it a second time. The old parent is now told before the new one adopts it.

  • Fixed SDLComponentNative's destructor resurrecting the Desktop singleton during shutdown: unregistering a window called Desktop::getInstance(), which constructs a new desktop when teardown has already destroyed the old one, tripping the check at the end of DeletedAtShutdown::deleteAll(). It now unregisters through getInstanceWithoutCreating() and skips it when the desktop is already gone.

  • DeletedAtShutdown now asserts at the point a new object is created while deleteAll() is deleting the others, instead of failing anonymously once the pass has finished and the stack naming the culprit has been unwound.

  • New ComponentNative::RepaintMode (ComponentNative::Options::withRepaintMode) selects how a window turns its accumulated dirty rectangles into repaint work. The default, RepaintMode::disjointRegions, repaints each dirty rectangle in isolation: a parent shared by several dirty rectangles is painted once, clipped to those rectangles, so components lying between two distant dirty rectangles are no longer repainted. RepaintMode::boundingBox keeps the previous behaviour of collapsing every dirty rectangle into one bounding box (repainting everything in between) and remains available as a fallback.

  • Transformed components are no longer clipped to their untransformed bounds. Component::internalPaint() now maps its local bounds through the transform accumulated from the component and its ancestors to the top level component before intersecting them with the accumulated dirty rectangles, so the redraw area of a scaled, rotated, or sheared component matches the area it actually covers and parts of it falling outside the untransformed bounds are painted again.

  • A paint() that throws no longer leaves the graphics frame open. SDLComponentNative::renderFrame() called context->begin() inside its render lambda but context->end() only after that lambda returned, so an exception from user paint code skipped the end()/tick() pair and the next frame began on a context that had never been flushed. The pair now runs from a scope guard

  • GUI: multitouch input on any platform with touch hardware (mobile, Emscripten in a mobile browser, and desktop touchscreens). Every finger is delivered through the existing mouse callbacks (mouseDown / mouseDrag / mouseUp, left button held): the first finger behaves exactly like a mouse, each additional finger arrives in parallel with a stable, dense finger index exposed as MouseEvent::isTouch() / MouseEvent::getTouchIndex(), plus touch pressure as MouseEvent::getPressure() (0.0-1.0). Each finger is hit-tested independently, keeps its index for the whole contact, and has its own double-click detection. SDL's synthetic touch-to-mouse events are disabled so every touch is delivered exactly once.

  • New "Touch Trails" example in examples/graphics: draws one hue-spaced colored trail per pointer (each finger gets its own color via MouseEvent::getTouchIndex(), with a mouse as a single pointer on desktop). Points carry an age in fade-timer ticks - not the wall clock - and their size and opacity decay with that age, so late frames can't make trails jump or flicker; the fade timer also runs while a pointer is down, dissolving the trail until only the pressed point remains, and once the finger lifts that point fades out too.

  • ApplicationTheme now exposes setDefaultMonospaceFont() / getDefaultMonospaceFont(), mirroring the existing default and icon font APIs. The default theme populates it from the embedded (or system) monospace font.

  • The default theme now embeds JetBrains Mono Variable (SIL OFL) as its monospace font when YUP_EMBED_DEFAULT_THEME_TEXT_MONOSPACE_FONT = 1 (forced on Emscripten), falling back to the system monospace font otherwise. The tools/embed_font.py regenerates the .inc byte arrays from any font file.

  • New CodeDocument (line-based text model with UndoManager-backed edits, positions, and incremental change notifications), SyntaxDefinition (JSON-driven language descriptions loaded from data/files or the built-in C++ / GLSL / Python definitions), CodeTokeniser (incremental per-line tokenizer with a line-state machine for multi-line constructs and lazy re-tokenization), and the CodeEditor component: syntax-highlighted editing, caret/selection with anchor semantics, clipboard, undo/redo, read-only, smart auto-indent, an optional line-number gutter with breakpoint markers, find/replace (find-all, next/previous with wrap, replace-one, replace-all in one undo step, match highlighting), bracket matching, and an optional minimap overview. Defaults to the theme's monospace font. See docs/ui/code-editor.md.

  • Added a built-in XML SyntaxDefinition (available as SyntaxDefinition::getBuiltIn ("xml") and matched for .xml, .svg, .html, .xaml and other markup extensions), with <!-- --> block comments, tag/attribute punctuation and <? ?> / <! / </ / /> operator highlighting. The CodeEditor demo now has a language dropdown to switch between the built-in C++ / GLSL / Python / XML definitions.

  • Added a built-in YDSP SyntaxDefinition (SyntaxDefinition::getBuiltIn ("ydsp"), matched for .ydsp), covering the language's keywords, primitive types and Faust-style composition operators (<:, :>, ->, ~, …). Used by the yup_dsp_jit "YDSP Synths" example's new Editor tab.

  • Fixed CodeDocument: newLineChars was default-constructed to an empty string instead of "\n", breaking getText(), getTextInRange(), and character-offset calculations for all multi-line documents; applyEdit() returned a wrong caret column for single-line insertions (omitted startIndex), making every subsequent undo call operate on an inverted range and silently no-op; removed the endsWithNewline special case that returned a pre-newline position and similarly broke undo for Enter at the beginning of a line or in the middle of a line.

  • Fixed CodeEditor: undo() and redo() now clamp caretPosition to the new document length and clear the selection after each operation, preventing an out-of-bounds caret after undo shrinks the document; replaceNext() now uses the position returned by replaceRange instead of selectionStart + replacement.length().

  • Fixed CodeTokeniser: cutting or deleting text that removes one or more lines left the token cache larger than the document and the forward-propagation stability check could declare a line whose content had shifted "unchanged", returning stale tokens (wrong syntax colors) for every line below the cut point. codeDocumentChanged now shrinks the cache to the new document line count and proactively marks all shifted lines dirty before the stability pass runs. The same problem existed in the other direction and was more visible in practice: inserting a line (pressing Enter, or a multi-line paste) grew the document but codeDocumentChanged had no branch for it at all, so every cached entry at or after the edit point kept referring to whatever used to be at that index - one or more lines off from where it actually was - and the stability check could decide a shifted-in line's state was "unchanged" and never mark it dirty, leaving it with stale, wrongly-sized tokens that fail to tile the line and fall back to unhighlighted plain text. Both directions are now handled the same way: resize to the new line count and mark everything from the edit point to the new end dirty, so misaligned cache entries are always discarded and recomputed from the live document text rather than reused.

  • Fixed CodeDocument::setText() freezing for several seconds on a large paste: its line-splitting helper indexed the (UTF-8-backed) input String by character position inside the split loop, and both operator[] and length() are O(n) for UTF-8, turning the split into O(n²). It now walks the text once with a CharPointer.

  • Fixed CodeEditor drawing selected/highlighted/caret text past the gutter and minimap when scrolled horizontally, since nothing clipped that content to the text area; the gutter's painted background also stopped 4px short of where the text area actually starts (disagreeing with the hit-test boundary used by mouseDown()), reading as misalignment on the left. The minimap overview now merges lines that map to less than one device pixel row into a single bar instead of issuing one fillRect() per source line on every paint regardless of visibility.

  • Fixed StyledText::update() calling Font::getPath() (a CoreText round-trip on Apple platforms) once per glyph occurrence instead of once per unique glyph; a glyph's outline is the same every time for a given font, so it's now cached and reused, cutting a measured 240ms of the 453ms spent reshaping text on a single keystroke.

  • Fixed CodeEditor reshaping (tokenizing, laying out and re-tessellating) the entire document on every single edit, making typing in a large file cost seconds per keystroke (measured: 2s for one backspace in a large file, mostly StyledText::update()). styledText now only ever holds the currently visible lines rather than the whole document; scrolling reshapes just the newly-visible range. Selection, search highlights, the caret, and Up/Down arrow navigation were adjusted to work correctly when their target is outside the currently-shaped range (falling back to an exact document-position computation rather than depending on styledText). Components that never call setSize()/setBounds() on their CodeEditor (as none of its unit tests do) keep shaping the whole document, since there's no meaningful "visible range" to restrict to without a real size.

  • CodeEditor now renders through a CodeEditorScheme (new code/yup_CodeEditorScheme.h): every color — background, gutter, caret, current line, selection, search highlight, breakpoint and the per-token syntax colors — is stored keyed by Identifier (CodeEditorScheme::setColor / getColor, string constants in CodeEditorScheme::ColorId) and switched with CodeEditor::setScheme. Built-in well-known schemes are provided via CodeEditorScheme::getBuiltIn: monokai, alabaster, oneDark, solarizedDark and solarizedLight. The editor's painting moved into the theme (Themes v1) as a registered CodeEditor component style, and a vertical auto-hide ScrollBar now appears when the document overflows the viewport. The CodeEditor demo gained a scheme dropdown.

  • Added responsive Rive Layout support to Artboard: per-node listeners (setNodeBoundsListener / clearNodeBoundsListener / clearAllNodeBoundsListeners / getNodeBounds) and attached components driven by layout reflows (attachComponentToNode / detachComponentFromNode / detachAllComponents, with NodeAttachmentOptions controlling whether the component fills the node's bounds or only tracks its position, follows its rotation via Component::setTransform, and — in track position mode — which component pivot point (pivot) is anchored to which point of the node's bounds (anchor), each accepting any Justification). Node listeners fire whenever the node's bounds or its on-screen orientation (rotation, scale, skew, mirroring) change and hand the callback a cached ArtboardNode handle — the same object Artboard::findNode returns — instead of allocating one per event, so the bounds/transform are read from the handle (which now also exposes getViewTransform(), the node's transform in the same component coordinates as getBounds()) and handles invalidate safely when the artboard is cleared or its file replaced. Reflows require the artboard's own layout (setFitting (std::nullopt)) or Fitting::fill, which propagate the component size into the Rive artboard (leaving those fit modes restores the artboard's authored size, and non-layout nodes report their real shape geometry instead of a unit rect).

  • Artboard::setFile() now takes an optional artboard name, loading that named artboard from the file (via rive::File::artboardNamed) instead of always loading the file's default artboard.

  • Added tools/rive_inspect.py, a standalone Rive (.riv) binary inspector with names, info, and tree commands (list named objects, dump object properties, and print the artboard hierarchy), resolving keys against the vendored rive generated headers without needing the runtime.

  • Added Rive ViewModel (data binding) support to the artboard module: ArtboardFile exposes the ViewModel schemas stored in a .riv file (getNumViewModels, getViewModelNames, getArtboardViewModel, createArtboardViewModelInstance) through new refcounted ArtboardViewModel (schema introspection: typed PropertyInfo per property with input/output flags and enum options, authored-instance names) and ArtboardViewModelInstance handles (typed and generic var property get/set by name or dotted path into nested viewmodels and lists, enum selection, trigger firing, list add/remove/swap/clear, and an optional setPropertyChangedCallback notified synchronously on value changes, including output bindings applied by the artboard). Artboard gained bindViewModelInstance / unbindViewModelInstance / getBoundViewModelInstance / getViewModelName, so a file's default instance can drive the artboard's data-bound properties and state machine transitions; a bound instance must come from the same ArtboardFile.

  • Artboard::setAllInputs() is now implemented (it was a documented no-op): it applies an Array<var> snapshot produced by getAllInputs() back through setInput(), matching entries by their "id" and ignoring unknown ones, so a snapshot can be saved and restored or moved between artboards. Triggers are stateless and deliberately carry no "value" in either direction.

  • ArtboardViewModelInstance paths may now terminate on a list index ("items.2"), which names the item's viewmodel instance and resolves through hasProperty() and getNestedInstance(); previously only value paths through an index ("items.2.quantity") worked and the two-segment form silently reported nothing.

  • ArtboardViewModelInstance structural list changes (addListItem, addListItemAt, removeListItem, swapListItems, clearListItems) now notify setPropertyChangedCallback with the list's own path and an empty var, as the callback's documentation already promised. Appending an item now extends the observer tree instead of rebuilding it wholesale, so filling a list is no longer O(N²) in observer allocations.

  • Artboard now auto-detaches a component attached with attachComponentToNode() when that component is destroyed, and attaching a component that already follows another node moves it rather than leaving it driven by both.

  • Artboard now re-derives an attached component's position when the component is resized. In trackPosition mode the component keeps its own size and that size is what the pivot is measured against, so resizing it after attaching used to leave the position stale until the node itself moved — which for a static layout is never. Size can now be set before or after attachComponentToNode() with the same result. In fillNode mode the node owns the size, so an owner-driven resize is put back.

  • Added docs/ui/artboard.md, covering file loading and asset resolution, layout/alignment, state machine inputs and events, node access and component attachment, and the ViewModel data-binding flow.

  • Artboard::getViewModelName(), ArtboardFile::getViewModelNames() and ArtboardViewModelInstance::hasProperty() are now const / no longer noexcept where they allocate; ArtboardViewModel and ArtboardViewModelInstance no longer type-erase their Rive pointers through void*, and Artboard / ArtboardFile gained the leak detector and non-copyable declarations the rest of the module already carried.

  • Fixed Component::getScreenPosition() and MouseEvent::getScreenPosition() adding the source component's own offset twice. getScreenPosition() computed localToScreen (getPosition()), but localToScreen() already adds getPosition(), so any nested component reported a screen position shifted by its own parent-local offset (and MouseEvent::getScreenPosition() inherited the same fault). Both now map through localToScreen() exactly once, so getScreenPosition() agrees with the origin of getScreenBounds().

  • GUI: drag-and-drop payloads can now carry arbitrary MIME data and a same-process native object. DragAndDropData stores an Array<ClipboardData> plus a var and offers withImage/getImage (PNG-encoded), the generic withMimeData/getMimeData/getMimeTypes/getAllMimeData, and withNativeObject/getNativeObject, with files and URIs sharing the text/uri-list MIME type. The class lives in modules/yup_gui/dragdrop/.

  • GUI: drag-and-drop targets are now an opt-in DragAndDropTarget mixin instead of five Component virtuals, so Component stays free of drag-and-drop and the whole feature is isolated in modules/yup_gui/dragdrop/. The SDL backend resolves targets with a dynamic_cast and bubbles from the deepest component under the cursor, preserving the previous enter/move/exit and drop semantics; DragAndDropSourceDetails carries the payload, source component, target-local position, allowed actions and suggested action. See docs/ui/component-drag-and-drop.md.

  • GUI: added DragAndDropTargetComponent, a concrete Component that is also a DragAndDropTarget, for the places that need one nameable type (factories, containers, the language bindings). The Python bindings are complete again: yup.DragAndDropData gained withImage/getImage/hasImage, withMimeData/getMimeData/getMimeTypes/hasMimeData and withNativeObject/getNativeObject/hasNativeObject, and yup.DragAndDropAction / yup.DragAndDropActions, yup.DragAndDropSourceDetails and yup.DragAndDropTargetComponent are bound - a Python subclass of the last overrides the target callbacks (or assigns the onItem* callables) and now actually receives drops, because it subclasses a single C++ type rather than deriving from Component and DragAndDropTarget separately (which pybind would give two unrelated C++ subobjects that the dynamic_cast dispatch could never find).

Audio GUI (yup_audio_gui)

  • SpectrogramComponent now keeps its waterfall history on the GPU: a precompiled .ysl shader bundle (embedded in yup_SpectrogramComponentShader.inc, built with the yup_shader_bundler host tool) drives a single fullscreen-triangle GpuRenderPass (see GpuPipeline) that scrolls the previous frame down by the pending rows and writes the new rows with the color map applied entirely on the GPU, uploading only the raw magnitudes as a uniform buffer - no per-paint CPU pixel upload, no GPU texture creation, and no 2D canvas flush (if the bundle cannot be compiled no waterfall is rendered). Pending FFT rows are always consumed (applied or dropped) so the update queue can never accumulate. The log-frequency → FFT-bin mapping is precomputed once per configuration instead of recomputed with pow/log per row, and the frequency grid (lines + labels) is cached in an offscreen canvas and only re-rendered when the frequency range or size changes. The component now requires a GPU render context (the CPU Image fallback was removed). The component's per-frame refreshDisplay hook processes pending FFT rows, and the history is presented at a fractional vertical offset that slides the newest row into place over one row period; that offset is pulled back with a fixed time constant instead of being clamped, so the motion never stalls at the edge of its window. The scroll speed is adjustable via the new setScrollSpeed() multiplier (1.0 = realtime, 0.0 = paused).
  • SpectrogramComponent's waterfall scroll is now independent of the frame rate. The component requests a repaint on every frame while the waterfall is live (through the existing refreshDisplay hook) instead of only when an FFT row arrives, so the sub-row offset above is actually rendered - previously the repaint cadence matched the row arrival cadence exactly, which made the display advance one whole row per repaint (one pixel, for a component whose height matches the history). A frame that produced more FFT rows than one GPU pass can write (a pass writes defaultSpectrogramMagnitudes / defaultSpectrogramWidth rows) now drains them over several passes instead of dropping the surplus, which is what happened whenever the display ran below the FFT row rate: at 30 fps half the history of a 2048/1024 configuration was silently lost. setScrollSpeed (0.0) now truly freezes the waterfall (pending rows are discarded rather than written, so the content no longer keeps sliding down while paused), and the animation stops requesting repaints once the analysis stalls instead of spinning on a frozen frame
  • SpectrogramComponent waterfall failures (shader bundle load, pipeline compile, and GPU pass encode/draw) are now reported via Logger::outputDebugString in all build configurations instead of silently dropping pending rows, and the waterfall texture's render resolution is exposed as the new defaultSpectrogramRenderWidth constant (2x the frequency-bin count - getSpectrogramImage() returns that full-resolution image).
  • Fixed SpectrumAnalyzerState never flagging FFT data as ready after a single bulk pushSamples(): the readiness check ran before the scoped FIFO write had committed (the AbstractFifo::ScopedWrite commits in its destructor), so isFFTDataReady() stayed false until a second push arrived. pushSample()/pushSamples() now commit the write before checking, so a pushed window is immediately available to SpectrogramComponent::refreshDisplay() instead of leaving the backlog untouched.
  • SpectrumAnalyzerComponent and SpectrogramComponent no longer snap every display point to its nearest FFT bin, which rendered identical levels (a flat staircase) for all the points sharing one bin. The new SpectrumBinMapping helper (displays/yup_SpectrumBinMapping.h) holds the shared log-frequency → fractional FFT bin mapping and evaluates levels continuously: the three bins surrounding a fractional position are parabolically interpolated in the amplitude decibel domain, evaluated at that position rather than at the vertex of the parabola, with a monotone linear fallback where the neighbours are not concave so a steep bin pair cannot undershoot. Display bands are aggregated over their fractional edges - peak for the peak/RMS level modes, sum (the levels integrated across the band width, in bin units) for powerDecibels and mean for powerSpectralDensity - so bins entering or leaving a band no longer step the displayed level. In the power modes this makes the band power an integral over the band's bandwidth instead of a sum over the integer bins its edges happen to touch
  • SpectrumAnalyzerComponent builds its spectrum outline once per pixel column, interpolating between the 512 smoothed display points, so the curve stays continuous at any component width and at HiDPI scales instead of following the 512-point polyline verbatim. The private computeSpectrumPath (Path, …) became createSpectrumPath (const Rectangle<float>&, bool) returning a Path, removing its reliance on Path sharing its rive::rcp<RiveRenderPath> between copies

Layout

The FlexBox and Grid containers landed in this cycle (they were previously listed under 1.0.0 by mistake) and have since been reworked against a browser-derived conformance corpus - tests/data/layout/capture.html renders each configuration with real CSS and tests/yup_gui/yup_FlexBoxParity.cpp / yup_GridParity.cpp replay the recorded rectangles.

  • Breaking: FlexItem::width, height and flexBasis now use -1 to mean auto instead of 0, matching GridItem and the min/max fields. A value of 0 is now a genuine zero size, so FlexItem (component, 0, 0) - which used to mean "auto on both axes" - is a zero-sized item; use FlexItem (component) instead. This is what makes flex: 1 1 0 (equal shares regardless of content, the most common flex idiom there is) expressible at all, via the new withFlexBasis (0)
  • FlexBox: added withFlexShrink() and withFlexBasis() builders; those two fields previously had no fluent setter
  • FlexBox: align-items: stretch - the default - now actually stretches in a single-line (noWrap) container. The line's cross size was only expanded for multi-line containers, so an item with no explicit cross size laid out 0px tall, which is why nested containers with unsized children collapsed. Per CSS, align-content does not apply to a single-line container at all and its one line spans the container's whole cross size; it does still apply to a wrapping container that happens to produce one line
  • FlexBox: align-items: stretch no longer overwrites an explicit cross size - height: 50px in a stretch row stays 50px and aligns to the cross start
  • FlexBox: justify-content no longer double-counts space already consumed by flex-grow. The free space was computed before flexible lengths were resolved and then reused as the flexEnd/center offset, pushing items outside the container; it is now recomputed after the flex pass, so a line whose items grow to fill it leaves justify-content nothing to shift by
  • FlexBox: flexible lengths are resolved with the CSS §9.7 freeze-and-loop instead of a single proportional pass. An item that hits its maxWidth while growing (or its minWidth while shrinking) now freezes and the space it could not take is redistributed over the remaining items, so a line no longer leaves a residual gap or overflow
  • FlexBox: align-content: spaceBetween now accounts for the preceding lines' cross sizes. Middle lines were positioned from the container's start with no accumulated offset, so with three or more lines they overlapped the first one. Every align-content mode now goes through one leading-offset/spacing computation
  • FlexBox: line breaking counts the gaps already consumed on the line. The accumulator tracked item sizes only while the free-space math counted gaps, so wrapped lines overflowed by roughly (n-2) × gap and broke one item too late. Breaking also now uses the hypothetical main size (the base size clamped by min/max) as CSS requires
  • FlexBox: spaceAround is no longer space-evenly on either axis. CSS puts half a share at each edge and a full share between (extra/2n and extra/n); both justifyContent and alignContent used extra/(n+1) everywhere. spaceBetween and spaceAround now also fall back to flex-start and center respectively when the content overflows, per the box alignment spec
  • FlexBox: reverse directions mirror the item's position within the line rather than the final rectangle, so marginLeft stays a left margin in rowReverse (and marginTop a top margin in columnReverse) instead of silently acting as its opposite. wrapReverse likewise flips the cross axis rather than only reversing the line order, so align-items resolves against the flipped axis
  • FlexBox: order is applied with a stable sort, so items sharing an order value keep their source order as CSS requires
  • FlexBox: baseline alignment no longer subtracts the leading cross margin twice, derives the default baseline from the item's clamped cross size, and grows the line when a baseline-aligned item needs more room than the tallest item alone. AlignItems::baseline in a column container now falls back to flex-start, matching browsers - the cross axis there is the inline axis, where a box with no text has no baseline to share
  • FlexBox: negative flexGrow, flexShrink and gap are asserted and clamped to 0 instead of being used as-is
  • FlexBox: removed calculateLayout(), which was declared but never defined or called, and the LineInfo struct it exposed, whose size fields were copied before being initialised
  • Grid: fractional (fr) tracks now subtract the gaps before dividing up the remaining space. Track positions advance by size + gap, so every gapped fr grid previously overflowed its container by exactly the total gap
  • Grid: placement follows the CSS order - all explicitly positioned items are recorded first, then items locked to a row pick a column within it, then the rest flows. An auto-placed item can no longer claim a cell that an explicitly placed item further down the list owns, and an item that pins only one axis keeps it (a single "either is unset" test previously discarded both)
  • Grid: item sizes and cell sizes are clamped at 0, so margins larger than the cell no longer produce an inverted rectangle
  • Grid: added GridItem::autoPlace as the named constant for the automatic-placement sentinel, and documented that YUP's grid lines are 0-based whereas CSS numbers them from 1
  • Grid: calculateTrackSizes moved out of the class (it never used this), and the auto-placement scan is bounded so a pathological span or position cannot allocate without limit
  • FlexBox/Grid: performLayout no longer rounds to whole pixels. Component::setBounds takes floats, and rounding position and size independently made adjacent items disagree about the edge they share
  • Documented that Grid::TrackInfo::auto_() is a fixed autoRows/autoColumns size rather than CSS's content-driven auto, and that autoRows/autoColumns also size auto tracks inside a template, not just implicit ones. Content-driven sizing needs a measurement hook on Component that does not exist yet; the one place the layout engines ask an item for its content size is now a single documented function in yup_FlexBox.cpp so it can be swapped without re-plumbing the algorithms

FlexBox feature parity

  • Added spaceEvenly to FlexBox::JustifyContent and AlignContent, which is an equal share between items and at both edges - distinct from spaceAround's half-share edges
  • Added independent rowGap and columnGap. gap is now a shorthand used only for whichever of the two is left at -1; in a row container the columns separate items and the rows separate lines, and in a column container it is the other way round
  • Added container paddingLeft / paddingRight / paddingTop / paddingBottom with setPadding() shorthands, which shrink the content box every item is laid out in. Padding larger than the target area clamps instead of inverting it
  • Added CSS margin: auto via FlexItem::marginLeftAuto and friends (withAutoMargins()). An auto margin absorbs its share of the free space on its axis before justifyContent is consulted, so marginLeftAuto pushes an item to the far end of a toolbar and a pair of them centers it. On the cross axis an auto margin positions the item within its line and suppresses stretching
  • Added FlexItem::flexBasisPercent (withFlexBasisPercent()), resolved against the container's main-axis size and taking priority over flexBasis

Grid feature parity

  • Breaking: Grid::TrackInfo is now a min/max pair (minimum / maximum, each a SizingFunction) instead of the three parallel pixelSize / fraction / isAuto fields. All construction still goes through the factory functions, which are unchanged, so only code that read those fields directly is affected
  • Added Grid::TrackInfo::minmax(), percent() and fitContent(). minmax (px (100), fr (1)) is the combination fr alone cannot express: take a share of the leftover space but never drop below 100px. Percentages resolve against the container's full size, not against what is left after the gaps
  • Track sizing now follows CSS §12.4-12.7: tracks start at their minimum, leftover space grows them towards their maximum, and the fractional tracks then divide up what remains with the same freeze-and-loop as flex-grow - a track whose share would land below its own minimum freezes there and the rest is redistributed
  • Added Grid::repeat (count, track) and Grid::repeatToFill (track, size, gap, defaultSize), the latter being CSS's repeat(auto-fill, ...). Note that auto-fit is deliberately not offered: it differs from auto-fill only by collapsing tracks that end up with no items, which needs to know each track's contents, so offering both would promise a difference that is not there
  • Added Grid::autoFlow with row, column, rowDense and columnDense. Placement was previously row-only and sparse-only; a dense flow restarts the cursor for each item so it can backfill the holes a larger item left behind
  • Added Grid::justifyContent and alignContent, so a fixed-track grid can finally be centered (or spaced) inside a larger box instead of always leaving the slack at the bottom right
  • Added Grid::setTemplateAreas(), mirroring CSS grid-template-areas, with GridItem::withArea(). The call returns a Result and rejects ragged rows and non-rectangular areas rather than laying out something surprising; a . cell belongs to no area and stays available to auto-placement
  • Added named grid lines: Grid::setColumnLineName() / setRowLineName() with GridItem::withColumnStart() / withRowStart(). Names resolve to YUP's 0-based indices at the point they are declared, so the 1-based numbering CSS uses never reaches the placement code
  • Added baseline to Grid::AlignItems and GridItem::AlignSelf, matching FlexBox. Items sharing a row align on a synthesized baseline - the item's bottom edge, which is also what a browser uses for a box with no text
  • Added a gap shorthand to Grid, so rowGap / columnGap now default to -1 meaning "use the shorthand", matching the new FlexBox fields
  • New docs/ui/component-layout.md concept guide covering both containers, the -1-means-auto convention, and a table of every deliberate divergence from CSS. docs/ui/component-basics.md no longer claims YUP has no layout manager
  • LayoutDistribution (new layout/yup_LayoutDistribution.h) is now the single definition of what each alignment mode means, shared by FlexBox's justifyContent / alignContent and Grid's. Four separate copies of that switch are what let spaceAround be implemented as space-evenly in some of them and correctly in none

Python bindings

  • FlexBox, FlexItem, Grid, GridItem, their enums, Grid::TrackInfo (factory-only, as in C++) and Grid::repeat() / repeatToFill() are now exposed to Python, with the layout containers' items and templateColumns / templateRows bound as live arrays. New python/tests/test_yup_gui/ covers both containers against the same expectations the C++ tests assert. Note that an item only stores a raw component pointer, so a component must be kept alive by the caller for as long as the item referencing it is used

Shading

  • New GLSL→WGSL direct transpiler in yup_shading: parses preprocessed GLSL 4.50, lowers GLSL constructs to WGSL equivalents, and emits WGSL 1.0 source. Supports vertex/fragment/compute stages with full builtin mapping, combined sampler splitting, entry-point IO wrapping, and binding assignment matching glslang's SPIR-V assignment 1:1. Does not require SPIR-V or spirv_cross for code generation. Integrated into ShaderTranspiler, ShaderCache, and ShaderBundleCompiler. WGSL variants are supported in YSLB bundles via the shader_bundler tool and yup_add_shader_bundle() CMake helper.
  • New GpuPipeline class (rhi/yup_GpuPipeline.h): an immutable compiled render pipeline (vertex + fragment shaders plus fixed pipeline state). compile(ctx, vs, fs, GpuPipelineOptions), compileFromBundle(ctx, ShaderBundle, GpuPipelineOptions), and (when YUP_ENABLE_SHADER_TRANSPILER = 1) compileFromGlsl(ctx, vertexGlsl, fragmentGlsl, GpuPipelineOptions) all return ResultValue<GpuPipeline::Ptr>. Pipelines carry all the backend-agnostic mirror enums/structs (GpuVertexFormat, GpuPipelineOptions, GpuColorTarget, GpuDepthStencilState, …).
  • New GpuFrame class (rhi/yup_GpuFrame.h): move-only RAII GPU frame scope (GpuFrame::begin(ctx) → submit() → waitForGPU()). Owns the transient GPU resource pools (uniform buffers, texture views, samplers) created while encoding its passes.
  • New GpuRenderPass class (rhi/yup_GpuRenderPass.h): move-only transient render-pass encoder targeting a GpuCanvas. Holds the mutable binding state (setPipeline, setTexture, setUniformBuffer, setVertexBuffer, setIndexBuffer) and encodes draws (draw, drawIndexed, finish).
  • New GpuPipelineCache class (rhi/yup_GpuPipelineCache.h): thread-safe compile-or-fetch cache for GpuPipeline keyed by a deterministic SHA1 of the selected native shader sources, entry points, pipeline options, and graphics API. LRU eviction with a configurable entry limit, mirroring ShaderCache.
  • New GpuBuffer class (rhi/yup_GpuBuffer.h): reference-counted GPU buffer handle wrapping a backend-native GPU buffer. GpuBuffer::create(ctx, GpuBufferType, data, byteSize) uploads immutable vertex/index/uniform data for use with GpuRenderPass.
  • Image::fromTexture(GpuTexture::Ptr): creates an Image wrapping an existing GPU texture (no CPU round-trip). Suitable for Graphics::drawImage().
  • Graphics::drawTexture(GpuTexture::Ptr, Rectangle<float>): draws a GPU texture directly without materialising an Image, avoiding CPU-side ImagePixelData allocation.
  • GpuRenderPass no longer creates a sampler and a uniform buffer per draw. The linear/clamp-to-edge samplers that fill a layout's sampler bindings are created once when the GpuPipeline is compiled, and uniform buffers come from a size-bucketed pool on the GpuDevice that recycles them when a frame reports GPU completion - so a steady-state workload stops allocating GPU objects after its first frames. GpuFrame stays stack RAII; nothing changes for callers.
  • Fixed the GLSL→WGSL transpiler rejecting comma-separated members in a struct or interface block (uniform Params { float s, r, rx, ry; }), which failed with Expected ';'. Each declarator now becomes its own member and binds its own array specifiers.

Shader Compiler (#126 and #130)

  • New glslang (thirdparty/glslang), SPIRV-Cross (thirdparty/spirv_cross) and SPIRV-Tools (thirdparty/spirv_tools) for shader reflection and cross-compilation (GLSL, ESSL, HLSL, MSL).
  • New yup_shading module for cross platform shader handling.
  • New ShaderBundle class (shading/yup_ShaderBundle.h): RIFF binary format (.ysl) that stores original source, per-stage SPIR-V, all transpiled variants (GLSL/ESSL/HLSL/MSL), and full ShaderReflection data. Persists to / loads from OutputStream, File, and MemoryBlock via saveToStream / loadFromStream and friends. Lookup by stage + language via findShader().
  • New ShaderBundleCompiler class (shading/yup_ShaderBundleCompiler.h): drives ShaderTranspiler to compile + transpile multiple stage/language combinations in one call and returns a fully-populated ShaderBundle. Accepts a ShaderBundleCompileRequest with per-stage ShaderBundleEntry items (stage, target languages, TranspileOptions).
  • New BinaryOutputArchive / BinaryInputArchive pair (yup_core/serialisation/yup_BinaryArchive.h): binary stream archives that plug into the SerialisationTraits system; used internally by ShaderBundle to serialise ShaderReflection data into REFL RIFF chunks.
  • New standalone yup_shader_bundler console tool (cmake/tools/shader_bundler): takes a .vert and .frag GLSL (v450 Vulkan dialect) pair on disk and produces a single .ysl bundle containing transpiled variants for all target languages (GLSL/ESSL/HLSL/MSL).
  • New yup_add_shader_bundle() CMake helper (cmake/yup_shader_bundler.cmake): builds the yup_shader_bundler tool for the host once (cached in the global property YUP_SHADER_BUNDLER_EXECUTABLE), runs it at configure time to generate the .ysl, and embeds it into a linkable object library via yup_add_embedded_binary_resources. Works even when the outer build is cross-compiling, since the tool is built in its own host binary tree without forwarding the cross toolchain. Accepts an OPTIONS argument that forwards arbitrary extra flags verbatim to yup_shader_bundler (e.g. --spirv-opt, --target-langs, -DNAME=VALUE, -I<dir>).

DSP (yup_dsp_jit)

  • New yup_dsp_jit module (modules/yup_dsp_jit): YDSP, a realtime JIT-compiled audio DSP language compiled to native machine code via asmjit_library (x86-64 and AArch64). Full compiler pipeline (lexer, parser, type system with realtime-safety enforcement, optimiser, AsmJit backend) plus a zero-allocation realtime runtime (YdspCompiler, YdspAudioGraph). Supports Faust-style composition algebra and Cmajor-style processor/graph definitions, sample and block processing, history state, parameters and meters, and sidechain/scratch buffers.
  • YdspCompiler::compile() gains an optional importBasePath argument: import directives in a patch now resolve relative to that directory (the patch's folder) instead of the process working directory, and nested imports inside an imported file resolve against that file's own folder. This makes multi-file patches loadable from disk; the "YDSP Synths" demo passes each patch's path and ships five importable effect processors in data/synths/fx/ (Delay, Compressor, Reverb, Distortion, Chorus), one wired into the graph of each demo synth.
  • Closed a set of silent-failure gaps found by an audit against the language spec: min/max/clamp/abs/sign gain dedicated integer opcodes (minI/maxI/clampI/absI/signI, branchless on both asmjit backends, compare+select on wasm) instead of only accepting float operands; endpoint annotations ([[ key: value ]]) now go through a whitelist that warns on an unrecognized key, matching every other annotation scope, and unit/step/style are plumbed all the way to YdspParameterInfo and the YDSP Synths demo's slider setup; stream[N] with N != 1 is now a compile error instead of silently yielding mono; buf[i] += x and s.field += x (and every other compound-assignment operator, including the previously-missing /=/%=) now desugar correctly via a deep-cloned target instead of failing with "Unknown symbol ''"; the lexer accepts .5, 1., 0x1F, 0b1010 and 1_000 literal forms and \n/\t/\"/\\ string escapes; automating a non-float32 parameter is now counted in getDroppedEventCount() instead of silently discarded; and the function inliner has a re-entrancy guard, so recursion the analyzer's own check misses now fails with a diagnostic instead of overflowing the host process stack.
  • The one-per-test YdspCompiler/YdspAudioGraph recompiles in yup_YdspGraphTests.cpp's YdspElectricPianoTests fixture and yup_YdspExamplePatchTests.cpp's two shipped-patch sweeps are replaced with a compile-once cache (yup_YdspTestPatches.h's cachedPatch/restoreFreshState) and a single merged sweep, cutting redundant JIT compiles in the test suite.

AI (yup_ai)

  • New yup_ai module (modules/yup_ai): LLM client and AI integration classes depending on yup_core and yup_events.

LLM

  • LLMClient (yup_LLMClient.h): abstract base for chat-completion backends with complete() and completeStreaming() methods, tool-call loop support via runToolLoop(), and structured output via LLMSchema JSON Schema or GBNF grammars.
  • LLMHttpClient (yup_LLMHttpClient.h): HTTP transport for LLMClient with retry and timeout logic, handling streaming SSE and non-streaming JSON responses.
  • LLMClientFactory (yup_LLMClientFactory.h): creates the correct LLMHttpClient subclass from LLMClient::Options::provider, with convenience factories for each provider.
  • LLMMessage (yup_LLMMessage.h): chat message with four roles (system, user, assistant, tool), optional tool calls, and serialisation to/from OpenAI ChatML JSON.
  • LLMResponse (yup_LLMResponse.h): parsed completion response with choices, token usage, tool-call extraction, streaming chunk accumulation, and error handling.
  • LLMTool (yup_LLMTool.h): callable function descriptor with JSON Schema parameters and a local handler, serialised to OpenAI function-calling format.
  • LLMToolRegistry (yup_LLMToolRegistry.h): thread-safe registry for LLMTool instances with snapshot, lookup, dispatch, and tools-array serialisation.
  • LLMSchema (yup_LLMSchema.h): fluent builder for JSON Schema objects (string, number, integer, boolean, array, object, oneOf) used in structured-output requests across all providers.

LLM Providers

  • LLMOpenAIChatClient (yup_LLMOpenAIChatClient.h): OpenAI Chat Completions API - also compatible with Ollama, DeepSeek, OpenRouter, and llama-server.
  • LLMOpenAIResponsesClient (yup_LLMOpenAIResponsesClient.h): OpenAI Responses API (GPT-5+, reasoning models).
  • LLMAnthropicClient (yup_LLMAnthropicClient.h): Anthropic Messages API (Claude models).
  • LLMGeminiClient (yup_LLMGeminiClient.h): Google Gemini generateContent API.

Embeddings

  • EmbeddingModel (yup_EmbeddingModel.h): OpenAI-compatible HTTP embedding model with embed() / embedBatch() and cosineSimilarity() helper.

MCP (Model Context Protocol)

  • MCPTypes (yup_MCPTypes.h): JSON-RPC 2.0 request/response/error types, MCP capability flags, tool and resource definitions with toVar / fromVar serialisation.
  • MCPTransport (yup_MCPTransport.h): abstract transport interface for JSON-RPC messages (stdio, HTTP/SSE, sockets, in-process).
  • MCPClient (yup_MCPClient.h): synchronous MCP client with initialize() handshake, listTools() / callTool(), listResources() / readResource(), and tool-import bridge registerToolsWith().
  • MCPServer (yup_MCPServer.h): MCP server exposing local YUP tools and resources over a transport, with registerTool() / registerResource(), start() / stop(), and placeholder startStdio() / startHttp().

Python Bindings

  • Python bindings for yup_ai (modules/yup_python/bindings/yup_YupAi_bindings.cpp): exposes LLM client, provider, messages, tools, responses, MCP types, client, and server to Python via pybind11.

Python Bindings (yup_python)

  • ComponentEffect.displayToContent / contentToDisplay and Component.getChildPointFromLocal / getLocalPointFromChild are overridable from Python; the point mapping helpers, setManuallyComposited() and renderToTexture() are bound.

  • GpuCanvas.beginDraw() accepts an optional scale, and Component.renderToTexture() an optional scale.

  • The new Component callbacks are overridable from Python: hitTest, childBoundsChanged, parentSizeChanged, indexInParentChildrenChanged, focusOfChildComponentChanged, keyStateChanged and modifierKeysChanged. hitTest is also callable directly, and the new yup.FocusChangeType enum is bound alongside the extra cause argument on ComponentNative.setFocusedComponent()

  • Bound the rest of Component's public surface: getSafeAreaBounds(), setPaintProfilingDisabled() / isPaintProfilingDisabled(), setMetric() / getMetric() / findMetric(), setCachedToTexture() / isCachedToTexture(), setComponentEffect() / getComponentEffect(), addComponentListener() / removeComponentListener() and snapshotToImage() / snapshotToTexture(). The effect and listener methods needed somewhere for their argument to come from, so ComponentEffect (whose apply() is the subclass point), ComponentListener and ComponentPaintMetrics are bound too, and addComponentListener() keeps the Python listener alive - the C++ listener list only holds weak references. PyComponent also gained the opacityChanged override it was missing, the one Component virtual no Python subclass could override

  • The drag-and-drop payload is bound as yup.DragAndDropData (withFiles() / withText() / withUris() builders, getFiles() / getText() / getUris() getters and hasFiles() / hasText() / hasUris() / isEmpty() predicates), and the five Component hooks that carry it - isInterestedInDrag(), itemsDropped(), itemDragEnter(), itemDragMove() and itemDragExit() - are callable from Python. A Python Component subclass could already override those hooks, but no drop could ever reach them: the payload had no Python type, so handing it back to the override threw the moment the platform delivered a drag. withFiles() takes any iterable of File, because the Array<File> binding's name is derived from the typeid at runtime and is not something to require of callers

  • yup.GpuCanvas.create() now accepts the three-argument form create(ctx, width, height). pybind11 does not see C++ default arguments, so the bound signature demanded the optional clearColor too and every call raised TypeError: create(): incompatible function arguments.

  • yup.GpuCanvas.beginDraw() returns a borrowing wrapper of the canvas's Graphics instead of trying to copy it. Graphics is non-copyable, so pybind11's default copy return policy made every call raise RuntimeError: return_value_policy = copy, but type yup::Graphics is non-copyable!; the wrapper now borrows the object and keeps the canvas alive alongside it.

  • Widget style identifiers are now reachable from Python: Label's nested C++ Style struct is bound as yup.Label.Style, so label.setColor (yup.Label.Style.backgroundColorId, yup.Colors.darkblue) works instead of raising AttributeError: type object 'yup.Label' has no attribute 'backgroundColorId'. The struct holds only static members and is deliberately not constructible

  • yup.Color now converts implicitly to yup.GpuColor, so yup.Colors.black can be passed anywhere a GpuColor is expected (GpuRenderOptions(True, yup.Colors.black)) instead of raising a TypeError. C++ already converts implicitly - GpuColor's constructor accepts any type with float-component accessors - so the binding registers that constructor and the conversion that uses it

  • Justification::Flags members are now implicitly convertible to Justification, so yup.Justification.left can be passed straight to any API taking a Justification (Graphics.fillFittedText and friends) instead of raising a TypeError. Combining flags with | yields the underlying integer, so an int converts too

  • An exception raised on a background thread no longer stops the dispatch loop. paint() runs on the render thread, and PyErr_CheckSignals() does nothing off the main thread, so the message-thread signal timer is the only thing that can ever notice Ctrl+C - stopping the loop from the render thread took that timer down with it and left the process deaf to SIGINT for the rest of its life. Only the message thread stops the loop now; a failing paint() reports and keeps running, so Ctrl+C and the window's close button still work.

  • Bound ApplicationTheme: yup.ApplicationTheme.getGlobalTheme() plus getDefaultFont(), getDefaultIconFont() and getDefaultMonospaceFont(), so Python code can obtain theme fonts the same way C++ does. getGlobalTheme() raises when no theme is set instead of returning a null pointer

  • Exceptions raised in Python overrides that C++ calls back into (paint, timerCallback, messageCallback, ...) now reach the application's unhandledException() with the original type and traceback, rather than terminating the process. START_YUP_APPLICATION's catchExceptionsAndContinue now governs the no-override fallback: when true the dispatch loop absorbs the error and keeps running, when false it stops so control returns to the interpreter. Previously that path called std::terminate() unconditionally.

  • TestApplication (the yup.TestApplication context manager used by the Python test suite) now keeps a reference to the application it constructs. The only py::object was a local in the scope's constructor, so the application was destroyed as soon as construction finished and YUPApplicationBase::getInstance() was null for the entire test. Everything guarded by it silently did nothing — most visibly sendUnhandledException(), so exceptions raised in Python overrides never reached unhandledException(). The scope now also calls shutdownApp() on teardown, which it never did.

  • Python exception reporting no longer prints each error's location twice. Both Helpers::printPythonException and PyYUPApplication::unhandledException combined error_already_set::what() with traceback.print_tb(), but what() already renders the traceback itself, so every error appeared once in pybind's At: file(line): func form and again in Python's. Both now use traceback.print_exception(), which produces the single rendering the interpreter would.

  • Ctrl+C is now noticed promptly instead of only when the application next happens to run Python on the main thread (which for a GUI whose only Python code is paint() on the render thread meant the next window focus or mouse event). The dispatch loop runs with the GIL released and CPython only runs signal handlers on the main thread, so runApplication now starts a message-thread timer that calls PyErr_CheckSignals() every messageManagerGranularityMilliseconds — a parameter that was passed but never used. It is a timer rather than a sliced runDispatchLoopUntil() because on Apple platforms that overload does not invoke the event loop callback that pumps SDL.

  • Ctrl+C stops a Python application again. Now that the dispatch loops catch rather than letting the KeyboardInterrupt unwind out of runDispatchLoop(), unhandledException has to stop the loop itself — it previously only set caughtKeyboardInterrupt and returned, leaving the loop running. runApplication also no longer re-raises a signal it has already reported.

  • Exception reporting no longer raises from inside the handler when the traceback is null, which is the case for a KeyboardInterrupt delivered by PyErr_CheckSignals() at a C boundary — a null py::object cannot be passed to a Python call and threw cast_error.

  • PyYUPApplication::unhandledException no longer dereferences a null std::exception*. sendUnhandledException() passes null for a catch (...), which the newly added YUP_CATCH_EXCEPTION sites make reachable.

  • Destroying a Python Component that is on the desktop no longer deadlocks the process. Tearing the native peer down joins the render thread, which can be blocked acquiring the GIL inside paint() / refreshDisplay() - and Python holds the GIL while it destroys the component, so neither thread could proceed and the message thread wedged inside the SDL event pump, taking the window's close button, the dock icon and Ctrl+C with it. The PyComponent trampoline destructor now detaches from the desktop with the GIL released, and addToDesktop() / removeFromDesktop() are bound with a gil_scoped_release call guard.

  • unhandledException reporting no longer touches Python from the thread that threw. Reporting from the render thread meant acquiring the GIL with no abort path while the message thread could be holding it to join that same thread, and after an exception it ran once per frame - import traceback, format, print - which is what made the deadlock above reproducible. Off the message thread the reporting and the stop request are now marshalled to it as a single message, so a stop can never latch the loop shut before the report is dispatched, and the throwing thread touches no Python at all. reportUnhandledException is also noexcept throughout: it is always called from a YUP_CATCH_EXCEPTION handler running in a C or Objective-C frame, where a raising unhandledException override would previously escape and leave the platform event locks permanently held.

  • An exception on the render thread now stops the application instead of repeating forever. unhandledException returned early on any thread that was not the message thread, so with catchExceptionsAndContinue=False a failing paint() was reported on every frame and the app never quit; stopDispatchLoop() is safe to call from any thread and is now used.

  • Ctrl+C no longer throws across the platform timer's C frame. The signal-check timer captured the raised error and stops the loop in place; START_YUP_APPLICATION re-raises it as a normal KeyboardInterrupt after the application has shut down. START_YUP_APPLICATION also always shuts the application down now - previously a re-raised exception or a signal checked after the loop skipped shutdownApp() entirely.

  • PyYUPApplication::unhandledException imported __builtins__, which is not a module in sys.modules, so the branch handling a non-Python exception raised from inside the exception handler. It now uses PYBIND11_BUILTINS_MODULE.

  • The yup_rhi bindings now cover the surface the RHI grew: textures (GpuTexture.create / upload, GpuTextureDesc, GpuTextureViewDesc, GpuTextureDataDesc), samplers (GpuSampler, GpuSamplerDesc, GpuRenderPass.setSampler), the full render-pass API (setColorAttachment, setDepthStencilAttachment, setResolveTarget, setViewport, setScissorRect, setStencilReference, setBlendColor, GpuDepthStencilOptions), vertex layouts and multi-target pipeline options, compute (GpuComputePipeline, GpuComputePass, GpuWorkgroupSize), the GpuTarget texture and view overloads with readPixels(), and the GpuDevice capability probes and buffer create/update/read entry points

  • Fixed GpuPipeline.compile* being uncallable from Python: they return ResultValue<T>, which has no registered Python type, so every call raised at return-conversion time. They now unwrap and raise the compiler diagnostics as a Python exception. GpuTarget.beginRenderPass was not bound at all, and GraphicsContext had no class binding, so Graphics.getGraphicsContext() raised as well - between them, neither python/demos/gpu_triangle.py nor gpu_effects.py could run

  • GpuRenderPass.setUniformBuffer and the buffer entry points take any buffer-protocol object (bytes, bytearray, memoryview, numpy arrays) instead of only bytes, and size them in bytes rather than items

  • GpuTextureFormat exposes all 25 formats (5 were bound), and GpuVertexFormat all 17 (7 were bound). Added GpuColorWriteMask (composable with | and &), GpuFilter, GpuWrapMode, GpuTextureType, GpuTextureAspect, GpuTextureViewDimension, GpuBufferType.storage and the remaining blend factors

  • GpuFrame and GpuRenderPass keep the objects they borrow alive (py::keep_alive), so a pass can no longer outlive the frame whose pools it points into

  • GpuShaderSource is now exposed, now that its blob fields own their data: code, bindingMap and glFixup are bytes-in, bytes-out properties accepting any buffer-protocol object. GpuPipeline.compile and GpuComputePipeline.compile are bound alongside the existing compileFromGlsl, so Python can compile from native shader sources without going through the GLSL transpiler. GpuCanvas.beginDraw() gained an optional GpuFrameDescriptor argument (also newly bound, along with GpuDitherMode), giving Python control over msaa/dither/loadOp/clearColor for the offscreen 2D frame it opens - bound as two overloads rather than one defaulted argument, since registerYupGraphicsBindings runs before registerYupRhiBindings and a default value referencing GpuFrameDescriptor at bind time would throw at import

  • Fixed registerYupRhiBindings being called under YUP_MODULE_AVAILABLE_yup_graphics while its translation unit compiles under YUP_MODULE_AVAILABLE_yup_rhi, which silently dropped the RHI bindings in a graphics-less build

  • Added python/demos/gpu_cube.py: a textured, depth-tested spinning cube driving vertex and index buffers, an uploaded texture and a sampler entirely from Python

  • python/demos/gpu_cube.py no longer renders its cube inside out. Its face corners wound counter-clockwise seen from outside, but the RHI bakes a clip-space Y-flip into the vertex stage — the GL backend inverts glFrontFace and Vulkan relies on naga's ADJUST_COORDINATE_SPACE to compensate for it — so GpuCullMode.back culled the side facing the camera instead of the far side. The corner order now matches SpinningCubeDemo's kCubeVerts

  • Fixed windows never appearing when a YUP application is run from a Python interpreter on macOS. A bare interpreter is not a bundled app, so the process starts as a non-UI one; SDL would normally set the activation policy, but it only does so when it is the one to create NSApp, and YUP's MessageManager gets there first. START_YUP_APPLICATION now transforms the process to a foreground application up front. Deliberately not done in initialiseYup_Windowing(), so a plugin hosted in a DAW can never transform its host's process

  • Added ColorGradient bindings (plus its Type / Spread enums and nested ColorStop). The previous binding had been commented out when the C++ API replaced bool isRadial with a Type enum, which left Graphics.setFillColorGradient() / setStrokeColorGradient() bound but uncallable

  • Component::paintSubtree now clears its isRepainting flag through a scope guard. A paint() override that throws - which is how a Python error surfaces - previously skipped the reset, leaving the flag set permanently so every later repaint() of that component tripped an assertion pointing at the wrong cause

  • Bound the yup_core facilities that had no Python surface at all: Logger (with setCurrentLogger() accepting a Python logger object), FileLogger, DynamicLibrary, SHA1, CancelToken / CancelTokenSource, WaitableTimer, StringPool, TextDiff, DynamicObject, AbstractFifo / SingleThreadedAbstractFifo, Expression (plus its ExpressionScope), LocalisedStrings, YAML (plus FormatOptions / Spacing), IPAddress, MACAddress, NamedPipe, WebInputStream and the InputSource family

  • registerStatisticsAccumulator() lets a C++ statistics type be addressed from Python the way the other generic templates are, which is what makes yup.StatisticsAccumulator[float] resolve instead of raising. Only float is registered: Python has a single float type, so registering double as well would map both spellings onto the same key and overwrite the first rather than adding an overload

  • Bound Fitting, CubicBezier and Drawable in yup_graphics; KeyModifiers, KeyPress, MouseWheelData, ProgressBar and SwitchButton in yup_gui; and MessageBase, Message, CallbackMessage and MessageListener in yup_events. The two widgets needed trampolines for the same reason the rest do - paint() is the subclass point

  • Fixed ImageFormatManager::createReaderFor leaking the stream it opened. The manager opens the file and hands the reader an InputStream to own, but createReaderFor released the Python wrapper of that stream on the way in while the format it built had no way to adopt it, so every call leaked one FileInputStream. ImageFormatReader now has a constructor that takes the caller's stream with deleteSourceWhenDestroyed set, and the bindings hand a Python format the stream the manager opened, so the reader deletes it exactly as the C++ path does

  • Methods that take a std::unique_ptr<T> parameter are only callable with a Python-constructed T when T is registered with py::smart_holder: pybind11 3.x moves the object out of the wrapper and disowns it, which it cannot do for a unique_ptr holder. InputSource (with FileInputSource and URLInputSource), XmlElement and the whole InputStream hierarchy (FileInputStream, MemoryInputStream, BufferedInputStream, SubregionStream, GZIPDecompressorInputStream, WebInputStream) were migrated, so XmlDocument.setInputSource(), XmlElement.addChildElement() and ZipFile.Builder.addEntry() now transfer ownership instead of releasing a wrapper nothing accounted for. The trampolines for those hierarchies also carry trampoline_self_life_support, which is a smart_holder requirement - pybind11 rejects the combination at compile time otherwise

  • MessageBase, Message and CallbackMessage are registered on their natural ReferenceCountedObjectPtr holder instead of the default unique_ptr. MessageBase::post() and MessageListener::postMessage() manage the message themselves - the queue takes its own reference, and the failure path only deletes a message whose count is still zero - so the previous bindings had to hand ownership over through release() to stop Python deleting a message the queue still held. With the refcounted holder both methods bind directly, and a message Python still holds stays valid after posting

  • These methods now consume the Python object they are given: after addChildElement(), setInputSource(), addEntry() or ZipFile(stream), using the wrapper raises ValueError: ... Python instance was disowned. That is the ownership the C++ API documents - the container deletes what it was given - and it replaces the previous behaviour, which leaked the reference instead

  • The new surface is covered by additions to python/tests/, and docs/scripting/python-bindings-coverage.md records the per-module inventory of bound against declared API that these gaps were found from

  • Bound MouseEvent, MouseListener and TextInputTarget, and declared the MouseListener base of Component. No Python Component subclass could receive a mouse callback that carried an event: mouseMove(), mouseDrag(), mouseUp(), mouseDoubleClick() and mouseWheel() were already routed to overrides, but yup.MouseEvent did not exist, so the moment the platform delivered one, the conversion back to Python threw. Component::addMouseListener() now also keeps the Python listener alive, matching addComponentListener(), and the wheel trampoline looks its override up under the C++ name mouseWheel - it asked for mouseWheelMove, a name nothing in the tree defined, so a wheel override written the obvious way was silently never called

  • Bound the yup_gui input and widget types the coverage page still listed as missing: ScrollBar, ListBox, ListBoxModel, ListBoxItem, ComboBox and TextEditor, each with its nested enums (ScrollBar::Orientation/VisibilityMode, ListBox::Orientation/SelectionMode, ListBoxItem::IconPosition) and its nested Style struct exposed for its theme identifiers the way yup.Label.Style already was. ListBoxModel is the subclass point for a Python list model and is dispatched through a trampoline; setModel() pins the model, which the ListBox never owns. TextEditor needed its own trampoline for getTextInputRect(), the TextInputTarget virtual it implements

  • ListBoxItem::setIconDrawable()/getIconDrawable() and ListBoxModel::refreshComponentForRow() are deliberately not bound: the first pair traffic in std::shared_ptr<Drawable> while Drawable uses pybind11's default unique_ptr holder, and the second hands the ListBox ownership of the component it returns, which cannot be taken away from a Python-owned instance without inviting a double free. setIcon() and paintListBoxItem()/getRowText()/getRowIcon() cover the same ground; both exclusions are recorded on the coverage page

  • Bound AudioIODeviceType and AudioDeviceManager::getAvailableDeviceTypes(), which python/demos/audio_device.py needs to list what the machine offers — the demo raised AttributeError: 'yup.AudioDeviceManager' object has no attribute 'getAvailableDeviceTypes'. The manager owns its device types, so getAvailableDeviceTypes() returns a fresh Python list whose entries borrow from it: a caller keeping a type keeps the manager alive alongside it, and createDevice() hands Python the device it creates with take_ownership, which is what the C++ contract asks for. AudioIODeviceType::Listener, addListener() and removeListener() stay unbound for want of a trampoline

  • PositionableAudioSource was bound without declaring its AudioSource base, so pybind11 treated the two as unrelated types: python/demos/audio_player.py died with TypeError: setSource(): incompatible function arguments ... (self: yup.AudioSourcePlayer, newSource: yup.AudioSource) ... Invoked with: ..., <yup.AudioTransportSource object>. An audit of all 289 py::class_ registrations against the module headers found this to be the only such omission (a scan that has to match the two-phase py::class_<...> classX (m, "X") form, since the single-phase declarations are all correct). AudioSourcePlayer::setSource() and AudioTransportSource::setSource() now also pin the source they are given with py::keep_alive: both C++ contracts say the object playing it does not own it, and nothing in Python could otherwise express that

  • AudioFormatReaderSource no longer takes deleteReaderWhenThisIsDeleted from Python, and no longer reads a freed reader. createReaderFor() hands Python a std::unique_ptr, and pybind11 cannot take a raw-pointer argument's ownership away from a live wrapper, so AudioFormatReaderSource(reader, True) — what python/demos/audio_player.py and audio_player_waveform.py both did — built a source that deleted a reader Python was about to delete too, and that read the reader after Python had dropped it. The source now always borrows and is pinned to the reader with py::keep_alive, so the reader outlives every use of it (C++ callers wanting the transfer still use the std::unique_ptr constructor). The symptom this fixes is not a crash: a dead reader reports a total length of 0, AudioTransportSource::hasStreamFinished() compares that with the read position, 0 >= 0, and getNextAudioBlock() clears the transport's playing flag on the first block, so the transport went silent and isPlaying() never became true

  • yup.AudioBuffer is subscriptable, so yup.AudioBuffer[float] resolves to yup.AudioBufferFloat the way yup.Rectangle[float] and yup.StatisticsAccumulator[float] already did — python/demos/audio_player_waveform.py failed with TypeError: type 'yup.AudioBufferFloat' is not subscriptable. The existing callable alias keeps working, so a subscription was added to it rather than replacing it with the type-keyed dictionary the other templated types expose: Python has one floating-point type, so such a dictionary could hold float alone and the double specialization stays reachable as AudioBufferDouble

  • python/demos/audio_player_waveform.py no longer glitches while it plays. Its paint() read the whole file through the same AudioFormatReader the audio thread was pulling, about 800 reads per frame: AudioFormatReader::read() re-seeks its stream and allocates on every call and the reader takes no lock, so the two threads moved the shared stream position out from under each other, and the audio thread competed for the allocator with a storm of render-thread reads. The peak envelope is now computed once, before playback starts, and painting only reads that list

  • python/demos/layout_flexgrid.py lays its panels out again. Every FlexItem was built with an explicit width of 0, and a literal 0 is a real zero size rather than "auto", so align-items: stretch skipped all five labels and the window showed nothing but the component's own black background. The sidebars were also added to a nested bodyFlex that never had performLayout() called on it, and self.content was added to both boxes - the nested row now runs over the band the column box leaves for the content, in a Rectangle[float] since Component::getLocalBounds() returns floats and RectangleInt rejects them

  • python/demos/layout_rectangles.py runs at all now. Its paint() had never executed past its first statement: w - 80 measured from float component bounds was fed to Rectangle[int], the eight .to<float>() conversions are the C++ spelling of the Python toFloat(), Graphics::drawText does not exist (fillFittedText (text, font, rect, justification) is the API, and the font comes from ApplicationTheme), Justification::centred became Justification::center, and the greys are darkgray/gray/lightgray - JUCE's British spellings were not carried over. It also carves its frame out of the window bounds with removeFrom* instead of subtracting from w and h: the subtraction went negative once the window was narrower than the 40px margin, and removeFrom* asserts on a negative extent (jlimit (0, extent, delta) in yup_Rectangle.h), so resizing the window down to zero width tripped it

  • Note for anyone else driving widgets from Python: laying widget text out goes through ApplicationTheme (ComboBox::updateDisplayText(), TextEditor's styled text and ListBoxItem::calculateLayout() all ask the theme for a font), and the global theme only exists while an application is initialised. Outside one those lookups dereference a null ReferenceCountedObjectPtr and take the process down, so a widget that carries text has to be created and used inside a running application - in the suite that is the juce_app fixture test_ApplicationTheme.py already used, which the new widget tests take as well

  • python/demos/matplotlib_integration.py is an actual port of popsicle's demo now, instead of a chart drawn out of YUP primitives. make_plot() / generate_plot_png() build the linear-regression figure in a child process - matplotlib is not thread safe, so it stays off the UI thread - and hand the PNG back over a multiprocessing.Queue; a 24Hz yup.Timer polls that queue, decodes the bytes with Image.loadFromData() into a child component that paints them, and fades that child in with setOpacity() while a star spinner turns behind it. Four YUP-for-JUCE substitutions were needed: the chart widget is a plain Component (YUP has no DrawableImage), the fade is driven from the timer (no Desktop::getAnimator()), fillAll() takes no color so the white fill is setFillColor() + fillAll(), and the child is added with addAndMakeVisible() because a YUP component starts out with isVisible() == false, where JUCE's starts visible. Two of those were only found by running it. The image child must not be opaque: Component::hasOpaqueChildCoveringArea() ignores the child's opacity, so an opaque child covering the parent makes internalPaint() skip the parent's paint() entirely and the window showed nothing but black - no white background, no spinner - until the chart arrived. And the timer is a plain ChartPoller(yup.Timer) holding the component rather than a second base of it, because a Python class deriving from two bound YUP classes (Component + Timer, as the original's MainContentComponent(juce.Component, juce.Timer) is) crashed with a bad this inside PyComponent's trampoline destructor when the window closed - the dealloc walk of such an instance runs pybind11's multiple-inheritance value_and_holder bookkeeping, and the class had been registered without the pair the pybind11 documentation asks of every trampoline. The Gui trampolines now derive from pybind11::trampoline_self_life_support and their classes are registered with py::smart_holder, which is what the Core and Graphics bindings had already been doing for their own trampolines. python/tests/test_yup_gui/test_MultipleInheritance.py covers it: the two-base case runs in a child interpreter, because the failure mode is a signal rather than an exception, and the test fails if the child is killed by one instead of reporting an error. Image.loadFromData() raises ValueError on a payload it cannot decode, where the ImageCache::getFromMemory() it replaces returned a null Image

Examples

  • The graphics synthesizer example's oscillator displays now draw the waveform the voices play with a yup_rhi fragment shader: a raymarched teal ridge landscape whose far ridge is one period of it, animating slowly while shown. The pipeline is compiled once and shared by both displays, and the vector display remains the fallback without a GPU.
  • The graphics synthesizer example's oscillator display can be drawn on freehand. The stroke is analyzed into 64 phase-preserving partials, which the PARTIALS view then edits.
  • Component3DDemo presents an interactive widget panel on a curved 3D surface; WidgetsDemo gains draggable corner handles applying an AffineTransform to the widgets panel; the Wave effect in ComponentEffectsDemo now maps input, with widgets inside the effected area.
  • Component3DDemo renders the panel texture and its 3D scene at the display scale.
  • Fixed the Python demo aborting on "Run Python!": PythonDemo is now bound with py::smart_holder, matching its Component base.
  • SpinningCubeDemo example (examples/graphics): rewritten to the new RHI shape - GpuFrame + GpuCanvas::beginDraw + GpuRenderPass for both the indexed cube draw and the separable two-pass blur (H+V sharing one GpuFrame), isGpuAvailable() capability probe, and live GLSL editing via GpuPipeline::compileFromGlsl. The default Lottie animation is now played back per-frame into an offscreen GpuCanvas (2D path) and sampled by the cube's fragment shader so the animation is texture-mapped onto every cube face.
  • WidgetsDemo example (examples/graphics/source/examples/Widgets.h): the placeholder image button is now a working ImageHitTestButton demonstrating Component::hitTest - it draws data/logo.png and samples the image's alpha at the hit point, so only the logo's opaque pixels are clickable and the transparent ones fall through to what is behind. The hover highlight goes through the same test, so moving the pointer over a transparent region inside the button's bounds drops it.
  • SpinningCubeDemo example (examples/graphics): rewritten to the new RHI shape — GpuFrame + GpuCanvas::beginDraw + GpuRenderPass for both the indexed cube draw and the separable two-pass blur (H+V sharing one GpuFrame), isGpuAvailable() capability probe, and live GLSL editing via GpuPipeline::compileFromGlsl. The default Lottie animation is now played back per-frame into an offscreen GpuCanvas (2D path) and sampled by the cube's fragment shader so the animation is texture-mapped onto every cube face.
  • AIDemo example (examples/graphics/source/examples/AI.h): interactive demo for all four LLM providers (OpenAI Chat, OpenAI Responses, Anthropic, Gemini) with model and API key configuration, system prompt editing, streaming and non-streaming completion, tool calling, MCP server integration, and embedded text generation.
  • ArtboardDemo example (examples/graphics/source/examples/Artboard.h): the loaded Rive artboard now queries a named node via Artboard::findNode (name, type, bounds shown in a status label) and attaches a rectangular marker component to it with Artboard::attachComponentToNode — a "Marker" combo switches between filling the node's bounds and tracking only its position, "Pivot" and "Anchor" combos choose which component point lands on which node point in track mode, and an "Apply transform" toggle rotates the marker with the node — and the marker follows the node on reflows and resizes.
  • New ArtboardLayoutDemo example (examples/graphics/source/examples/Artboard.h, registered as "Artboard Layout"): shares ArtboardDemoBase's controls with ArtboardDemo but loads data/layout-ui.riv and attaches a live, nested Artboard (playing the file's Keyboard artboard) to the keyboard_slot layout placeholder in the main Wireframe artboard, instead of a plain marker rectangle.
  • ArtboardDemo example: the displayed Rive file can now be replaced at runtime by dropping a .riv file onto the demo, which rebuilds the artboards from the dropped file while keeping the current fit, alignment and marker settings. The demo outlines itself while a .riv is dragged over it.
  • AudioExample example (examples/graphics/source/examples/Audio.h): reworked into a Vital-style instrument. Each oscillator gets a display that either draws the reconstructed waveform or edits its partials as draggable magnitude bars, writing sine coefficients the wavetable, sync and morphing backends all render; the drag publishes at most one generation bump per UI frame so it cannot queue an inverse FFT per mouse event across every sounding voice, and the reconstruction's peak is measured on the message thread and applied as a coefficient scale so an edited spectrum cannot exceed full scale. Unison adds up to five detuned, stereo-spread slots per oscillator, built from bare WavetableOscillator satellites rather than further SynthOscillator copies (which own four wavetable oscillators each once sync and morphing are counted) and offered on the wavetable algorithm alone, where they are exact. A DAHDSR envelope with a drawn curve replaces the single SmoothedValue fade, and the voice now renders stereo. The demo also plays the first available hardware MIDI input, collected through a MidiMessageCollector into the same buffer MidiKeyboardState reads, so external notes light up the drawn keys as well as sounding. Fixes the Detune control, which the UI wrote but the engine never read, so both oscillators always played the same frequency.

Build System

  • Added yup_add_bundled_resources() and a BUNDLE_RESOURCES argument on yup_standalone_app(), taking <file>@<relative-dest> pairs (same convention as PRELOAD_FILES) and placing them where File::getSpecialLocation (File::bundleDirectory) can find them at runtime: Resources/<relative-dest> in the app/plugin bundle on Apple, app/src/main/assets/<relative-dest> on Android, and preloaded into the Emscripten virtual filesystem
  • examples/graphics: the Rive artboard and Lottie demos now bundle their .riv/.lottie files via BUNDLE_RESOURCES (Android, iOS, Emscripten) and read them back through File::getSpecialLocation (File::bundleDirectory), instead of compiling them into the binary with yup_add_embedded_binary_resources()
  • justfile recipes now use per-platform build directories (build/mac, build/ios, build/android, build/emscripten, build/ninja, build/win), so switching platforms no longer requires just clean and preserves downloaded FetchContent dependencies per platform. The just build recipe gains a PLATFORM parameter (default mac).
  • yup_standalone_app gains a MAXIMUM_MEMORY Emscripten argument (-sMAXIMUM_MEMORY); when set it caps the heap that ALLOW_MEMORY_GROWTH may reach.
  • yup_tests wasm build: raised INITIAL_MEMORY to 256 MB, added MAXIMUM_MEMORY cap of 1 GB, and reduced STACK_SIZE to 1 MB to give the heap room for concurrent pthread stress tests; fixes RuntimeError: memory access out of bounds in CI.
  • Fetched third-party dependencies (SDL3, Perfetto, plugin SDKs) are now cloned shallowly (--depth 1) and skip network update checks on reconfigure, speeding up fresh configures and reconfigures. Shallow cloning is automatically disabled when a GIT_TAG is a commit hash.
  • Android: full 16 KB page size compatibility — generated Gradle projects bumped to AGP 8.5.2 / Gradle 8.7 (uncompressed native libraries are zip-aligned to 16 KB), jniLibs packaging made explicitly non-legacy, and ndkVersion pinned to r27c (overridable via NDK_VERSION), which ships a 16 KB-aligned libc++_shared.so. CI NDK updated to r27c accordingly. Application shared libraries were already linked with -Wl,-z,max-page-size=16384.
  • The vendored Rive runtime now builds with its scripting support enabled: the rive module declares WITH_RIVE_SCRIPTING=1 and RIVE_LUAU=1 and depends on the vendored luau (its freshly generated thirdparty/luau/luau.cpp amalgamates the VM sources) and libhydrogen (HYDRO_SIGN_VERIFY_ONLY=1) modules. Without the define, rive::File::read() has no importer for ScriptAsset/ShaderAsset, so an in-band FileAssetContents belonging to a script asset is routed to the previous asset's importer and trips assert(!m_content) in FileAssetImporter::onFileAssetContents — which is what made tests/data/rive/viewmodel-lab.riv abort on load (and would have silently overwritten a font/image's contents in a release build). Scripts still only reach the VM when their in-band signature verifies, so unverified script assets load as inert content.
  • The vendored Rive runtime no longer prints ScriptAsset doesn't have a generator function … once per scripted object per frame. A non-tools runtime registers only scripts whose in-band signature verified, so the generator lookup for an unverified script is expected to fail rather than worth a diagnostic. The message is dropped through a new patches entry on the rive dependency in tools/rive_update_manifest.json, so the next just rive_update reapplies it instead of the noise coming back.
  • Android: the vendored Rive runtime no longer defines RIVE_DESKTOP_GL on GLES platforms. It was defined whenever YUP_RIVE_USE_OPENGL was on and RIVE_WEBGL off, so Android took Rive's desktop GL path and gles3.hpp included glad's gles2.h. glad's glXxx function macros then aliased Rive's own (non-glad) extension entry points onto glad's identically named pointers, and the link failed with duplicate symbol: glad_glDrawArraysInstancedBaseInstanceEXT and eight siblings; had it linked, those pointers would also have been resolved by glad rather than by LoadAndValidateGLESExtensions(). Android now uses <GLES3/*.h> as intended.
  • Emscripten/WebGL: the OpenGL compute backend is no longer compiled, since WebGL (and WebGPU on the web) only reach GLES 3.0 and have no compute entry points. A new YUP_RHI_USE_GL_COMPUTE config guards the GL compute pipeline, pass, factories and dispatch sites, replacing the YUP_RIVE_USE_OPENGL || YUP_LINUX || YUP_ANDROID condition that held on every Emscripten build and failed with use of undeclared identifier 'GL_COMPUTE_SHADER', 'glDispatchCompute' and the GL_*_BARRIER_BIT enums.
  • Android CI: build_android.yml now tells android-actions/setup-android to install only platform-tools, avoiding its removed default tools package so the configure job can set up the SDK again

Testing

  • The YdspOptimizerTests suite (tests/yup_dsp_jit/yup_YdspOptimizerTests.cpp) is enabled and now exercises each optimizer pass - constant folding (including loop-carried induction registers, which must not fold), algebraic simplification, copy propagation, dead-code elimination and loop-invariant code motion - in addition to the existing IR-lowering and execution-report checks. The individual passes are exposed on YdspOptimizer so tests can drive them directly; the four-pass loop (constant folding, algebraic simplification, copy propagation, DCE) is wired into YdspOptimizer::runPasses, while LICM is exercised by its tests directly.
  • The AU and AUv3 wrapper tests no longer describe a stereo buffer list with a stack-allocated AudioBufferList. Only the first AudioBuffer is reserved inside the struct, so AUStateTests.RenderProducesOutput and the two AUv3BypassRenderTests render tests wrote past the object while filling mBuffers[1].mDataByteSize, aborting the suite under AddressSanitizer with a stack-buffer-overflow. All three now build their lists with a new tests/yup_audio_plugin_client/yup_TestAudioBufferList.h helper, which owns an allocation sized for the number of buffers asked for
  • The yup_events Python tests no longer depend on a single fixed-duration pump of the message loop. next(juce_app) runs the dispatch loop for 20ms and returns, but a dispatched callback still has to re-acquire the GIL before the Python side runs, so on a loaded machine it can land after the pump has already returned - test_MessageListener::test_construct_and_post failed this way on CI. The 22 call sites with a positive expectation now use a new pump_until(app, predicate) helper in python/tests/utilities.py, which pumps in short slices until the condition holds or a 5s timeout expires. The four sites that assert a negative after pumping ("still zero because it was cancelled") deliberately keep the fixed pump, since polling a negative predicate returns immediately and proves nothing
  • Component now befriends a single ComponentTestHelper<T> class template instead of accumulating one friend class per test suite; unit tests specialize it (e.g. ComponentTestHelper<Component>, ComponentTestHelper<ComponentEffect>) to reach private state.
  • Expanded the Rive viewmodel data-binding tests to run against tests/data/rive/responsive-sliders.riv, the data-binding showcase fixture: concrete schema assertions for both its ViewModels (Slider_instance, Main) and their authored instances, nested viewmodel / dotted-path value access and color round trips, cross-handle shared-value visibility, property-change callbacks reporting dotted paths, and full Artboard bind/advance/unbind coverage over the state-machine-driven "Main" artboard.
  • Added tests/data/rive/viewmodel-lab.riv and the coverage that goes with it. The fixture ships four ViewModels in a known order (Details, Row, Panel, Lab), authored instances for each, and a Lab schema declaring every property type the API models (string, number, boolean, color, custom enum, trigger, nested viewmodel, list, asset image, artboard reference). The new suites assert the schema contents (property order and types, the enum's Idle/Running/Failed options, the Row schema's input flags), the authored values of the Default and Preset instances - including both nested viewmodels, the three authored list rows and the per-instance list sizes - and the accessor contract for properties that carry no value (triggers, containers, symbol list indexes), plus bind/advance/unbind coverage of the "ViewModel Lab" artboard while writing values and mutating the row list. Every expected value is what tools/rive_inspect.py reports for the fixture, which is now also part of the cross-fixture schema sweep. Because the fixture ships ScriptAssets, the suites skip with an explicit message when Rive is built without WITH_RIVE_SCRIPTING. Running them found one defect: unbinding an instance whose list drives an ArtboardComponentList leaves the parent artboard's Yoga tree referencing the rows that the bulk teardown dropped, so the next layout pass walks freed nodes and crashes - that sequence is captured by DISABLED_UnbindingAndAdvancingKeepsTheArtboardUsable with a pointer to the cause rather than weakened to match the current behaviour.
  • Nine test files were globbed into the IDE project but never #included in their module's unity translation unit, so 141 tests had never been compiled or run since being written: yup_CodeEditorScheme.cpp, yup_ListBoxItem.cpp and yup_PaintProfileStats.cpp (yup_gui), yup_Memory.cpp and yup_TypeErasedObject.cpp (yup_core), yup_GraphicsContext.cpp and yup_ImageFormatMetadataExtended.cpp (yup_graphics), yup_AudioDeviceManagerWindow.cpp (yup_audio_gui) and yup_AudioPluginLV2Format.cpp (yup_audio_plugin_host). All are now wired up, which took three kinds of repair: two defined a file-scope helper that a sibling in the same unity build already defined (makeSample, loadFromBlock), so the newly enabled copies are renamed rather than the working ones; yup_ListBoxItem.cpp had drifted against the graphics API (PixelFormat is no longer nested in Image and has no ARGB, the Image constructor now takes (w, h, format), DrawablePath folded into Drawable, and String has no (count, char) constructor); and yup_AudioPluginLV2Format.cpp is now wrapped in #if YUP_AUDIO_PLUGIN_HOST_ENABLE_LV2, mirroring the guard the module puts around LV2Format, since the test target does not enable LV2. Wiring them up surfaced two genuine defects the dead tests had been written against: the PNG raw-chunk writer bug below, and ListBoxItem::setIcon being an unimplemented stub - the four tests depending on it are marked DISABLED_ with a pointer to the TODO rather than weakened to match the stub

Bug Fixes

  • YDSP state is now segmented as [scalars][arrays] and the kernel ABI carries both base pointers (state + new stateArrays in YdspKernelContext/YdspEventContext, with YdspCodegen::stateScalarSize reporting the split). Every scalar slot (including the ring write-pointers of the @ delay primitives) lives in the scalar segment head, and all arrays follow in their own segment. Previously, scalars were addressed at byte offsets that grew with array state (e.g. 50476 for a reverb kernel), which AArch64 could not encode as an LDR/STR immediate and failed to assemble with InvalidDisplacement (ldur w2, [x3, 50476]). Array state can now grow arbitrarily (delay lines, reverb rings) without pushing scalar slots out of range; an out-of-range materialization fallback remains as defense in depth. The optimiser's per-type array-base shift pass is gone (array bases are per-type element indices into the array segment).
  • SDL windowing: partial repaints now grow the dirty area by half a pixel before rounding it out to whole pixels. Rive applies rectangular clips as anti-aliased coverage rather than a pixel-exact scissor, so geometry touching a component's clip edge could bleed a tiny coverage into the adjacent pixel row; with a preserved render target that row was never redrawn and the bleed accumulated into a persistent line just outside components that repaint continuously (visible around the SpectrogramComponent at 1x scale). The extra border lets the parent repaint those pixels every frame.
  • AUv2 wrapper: an input bus that is not fed during a render cycle is now presented to the processor as a null-channel view. buildInputBusViews filled the per-bus channel pointers only for the channels it actually received, leaving the rest at whatever the previous render had stored there, so a sidechain input the host stopped feeding (inactive element or a failed PullInput) kept pointing at that element's stale audio instead of reading as silent - contradicting the comment on pullAuxiliaryInputElements and the AudioBusBufferView "null for an inactive or silent bus" contract. The input path now clears each bus's slots first, exactly as buildOutputBusViews already did for outputs; the AUv3 wrapper never had the problem because it maps input views onto its own scratch buffers
  • MessageManager on Apple platforms: runDispatchLoop() and runDispatchLoopUntil() now wrap their loop body in YUP_TRY/YUP_CATCH_EXCEPTION, matching the generic implementations in yup_MessageManager.cpp that #if ! (YUP_APPLE || YUP_WASM) compiles out on these platforms. The .mm replacements previously caught only NSException, a disjoint set from std::exception, so a C++ exception thrown by a message callback escaped the dispatch loop instead of reaching YUPApplicationBase::sendUnhandledException() — which made unhandledException() unreachable on macOS and iOS, and killed the application on the first failure.
  • The SDL render thread now routes exceptions from renderFrame() through YUP_CATCH_EXCEPTION as well. It never passes through a dispatch loop, so an exception escaping paint() reached Thread::threadEntryPoint(), which only asserts — rendering then stopped permanently with no diagnostic in release builds.
  • PNG: png/iCCP and png/cHRM raw chunks are now actually written. The iCCP branch in the writer was an empty if body with a "for simplicity, write as unknown chunk" comment that never wrote anything, and png/iCCP was excluded from the unknown-chunk loop; cHRM was collected but silently dropped by libpng, because on write libpng only emits unknown chunks whose name marks them safe-to-copy (a lowercase fourth letter) unless png_set_keep_unknown_chunks says otherwise - eXIf is safe-to-copy and survived, cHRM and iCCP are not and did not. Both are now registered with PNG_HANDLE_CHUNK_ALWAYS. The png_unknown_chunk is also value-initialised, since libpng copies all five name bytes and the terminator was left indeterminate
  • FlexBox: the gap is no longer applied after the last item on each line. It was added unconditionally after every item, so when items grew or shrank to fill the container exactly (e.g. flexGrow items with gap), the trailing gap pushed the final item past the container's main-axis edge and it got clipped (e.g. the last panel in each row of the Layout example).
  • ShaderTranspiler: GLSL ES fragment output now defaults to precision highp float; instead of SPIRV-Cross's precision mediump float; default. Desktop GLSL implies highp and glslang records no precision decorations for it, so SPIRV-Cross could only re-qualify declared variables; anything left to the fragment default — uniform block members and inlined expression intermediates — silently ran at mediump on OpenGL ES, corrupting fp32-exact math such as the fixed-point field codecs of the GPU fluid simulation demo (pixelated dye that never fades on Android). The ESSL default is now highp in both the emit and reflect paths.
  • Fixed data races in KMeterState: the per-channel getters (getPeakLevel/getAverageLevel/getPeakHoldLevel with an explicit channel index) read the plain ChannelState floats the audio thread updates, bypassing the atomics the aggregate (channel -1) path uses — those levels are now atomic, with the audio-side ballistics computing on locals and publishing once per block. The runtime configuration scalars (scale, meteringStandard, fall/hold times, over threshold/mode) are now also atomic since the processing thread reads them while setters run on the UI thread, and setMeteringStandard/setIntegrationTime/setPeakFallTime no longer mutate the loudness filters and level processors from the calling thread — they set a flag that processPendingAudio() applies on the processing thread that owns them. The class now uses std::atomic throughout (previously yup::Atomic)
  • StyledText caret bounds, hit-testing and selection rectangles now use line-relative glyph x positions computed with the same accumulation as drawing, instead of rive's paragraph-relative GlyphRun::xpos. Character positions were wrong on soft-wrapped lines (off by the width of all preceding text in the paragraph) and selection was drawn shifted on wrapped text; the caret at the first character of a wrapped line now lands on that line's left edge.
  • iOS applications now use the UIScene lifecycle, removing UIKit's legacy lifecycle warning and ensuring SDL windows are created for the connected scene.
  • Offscreen GPU rendering now supports recursive targets on Metal, OpenGL/GLES, and D3D11, so Lottie alpha/luma mattes, isolated-opacity layers, and cached precomps retain GPU compositing when rendered into an Image or GpuCanvas. Each RenderableTarget leases a Rive render context exclusively for its lifetime and returns it to the pool when destroyed. Repeated Lottie matte and precomp renders now reuse their canvases rather than allocating GPU textures each frame. Metal child targets allocate only their Rive render-canvas output texture; the CPU readback staging texture is created only when pixels are requested.
  • Fixed undefined offscreen contents when nesting pooled render targets on all GPU backends. Render context slots were recycled whenever no frame was currently active, so two long-lived targets could share one slot; once their frames nested - which happens as Lottie matte and precomp layers cross their in/out points and the nesting order changes between frames - the inner target skipped beginFrame and was then flushed against the outer target's frame descriptor.
  • Lottie: a matte layer no longer paints another matte layer's content. Drawing a matte result only queues a reference to its canvas texture, which the enclosing frame resolves at flush time, but the canvas lease was released as soon as the layer finished. Since every matte in a composition is sized to the same fitted rectangle, the pool handed the same canvas triple to the next matte layer, which overwrote the pixels already queued and left only the last matte visible (e.g. world_locations.json's four matted dots collapsed to one and its continent outlines disappeared; insta_camera.json lost its animated circles). Leases are now held until the composition render completes.
  • GpuFrame now waits for the GPU before releasing the texture views, uniform buffers and samplers it keeps alive for its encoded render passes. Those passes reference them by raw pointer, and submit() does not block, so letting a frame go out of scope freed them while the GPU was still reading — corrupting the pass output progressively, as the freed memory only starts being handed back out after the allocator has churned for a while (the growing magenta flashes in bell.json). waitForGPU() is now only needed explicitly when results are required before the end of the frame's scope, and is idempotent so waiting explicitly costs no more than one stall. Move-assignment drains the frame it replaces for the same reason.
  • AffineTransform::getScaleFactor() is now independent of rotation. It averaged the absolute values of the matrix diagonal and ignored the shear terms, so a rotated transform reported scale * cos(angle) — falling to zero at 90 degrees. It now measures the lengths of the transformed basis vectors. Lottie precomposition and matte canvases are sized from this value, so a layer under an animated rotation (e.g. bell.json, whose precomposition is parented to a rotating null) requested a different pixel size on every frame, reallocating its canvases mid-frame and flashing while a queued draw still referenced the previous ones.
  • Lottie: the matte canvas pool now replaces an idle slot of a different size instead of appending a new one. Nothing removed slots, so a layer whose on-screen size changed every frame added three canvases — each leasing a Rive render context — per frame, without bound.
  • GpuCanvas::create() takes a std::optional<Color> clearColor, defaulting to transparent black, and fills the new canvas with it so it is safe to sample before anything is drawn into it (pass std::nullopt to leave the contents undefined). The backing texture is allocated uninitialized and a 2D frame whose draw list ends up empty is not guaranteed to honour its loadAction=clear, so a canvas could previously be composited while still holding undefined GPU memory (the magenta flashes in bell.json, whose only content is one matted precomposition). The clear is issued through the new GpuDevice::clearOffscreen(), which encodes it with the backend's native API — a clear binds no pipeline, buffers or samplers, so it needs neither a render pass nor a submit/wait cycle.
  • Artboard::clear() and updateSceneFromFile() destroyed the rive artboard before the scene (StateMachineInstance) that references it, so destroying the scene called cleanupFocusTree() on freed memory (ASAN use-after-free). The scene is now reset before the artboard.
  • ArtboardViewModelInstance path resolution downcast a property to rive::ViewModelInstanceList without checking its type whenever a path segment was an index, so a path like "score.0" (where score is a number) asserted in debug builds and read a garbage vector in release ones. Reachable from every path-taking accessor. The four unchecked ViewModelPropertyEnum downcasts behind the enum accessors were guarded the same way.
  • Artboard only drained its state machine's reported events from mouseDrag(), so the documented onPropertyChanged / propertyChanged callbacks never fired for events reported by an advance, a pointer press or a transition — Rive clears the queue at the top of the next advance, and a resize or a bindViewModelInstance() (both of which advance) silently swallowed the frame's events. Every advance now drains through one helper, and every pointer handler drains after its pointer* call.
  • Artboard::notifyNodeBoundsChanged() iterated a member scratch array while invoking user callbacks that are allowed to call setLayout() / setAlignment(), which re-entered the function and cleared the storage the outer loop was walking (and the const String& its callee held). Nested passes are now skipped, and the loop tolerates a callback replacing or unloading the file.
  • ArtboardViewModelInstance invoked its property-changed callback in place, so a callback that re-armed or cleared itself (a one-shot listener) destroyed the closure that was executing — something the header explicitly documents as supported. The dispatch now runs a local copy.
  • Artboard::applyNodeAttachment() took its attachment record by reference straight out of the attachedComponents map, then read options.applyTransform from it after Component::setBounds(), which fires resized() synchronously. A resized() that detached the component (or attached it to another node) destroyed that map entry mid-call. The record is now taken by value.
  • Artboard::setFile (nullptr) dereferenced the null file instead of unloading, and the file-taking constructor did not call setOpaque (true) while the other one did.
  • ArtboardFile::AssetInfo::uniquePath was a File holding rive::FileAsset::uniqueFilename(), which is a bare file name ("logo-1234.png") and not an absolute path — so every asset-resolving load hit jassertfalse in File::parseAbsolutePath() and then silently resolved the name against the current working directory. Renamed to uniqueFilename and retyped as a String; resolve it against your own asset directory with File::getChildFile().
  • ArtboardFile::load() ignored the result of readIntoMemoryBlock(), so a stream that yielded nothing was reported as "Malformed artboard file" rather than as a read failure.
  • GpuCanvas::beginDraw() now drops the target's cached GpuTexture wrap, as its documentation already claimed. The wrap memoizes the Rive texture handle it resolved, so a pooled canvas reused across frames kept handing out the handle resolved on the frame it was first sampled.
  • Lottie: a failed matte composite no longer blits undefined GPU memory over the matted layer. The result canvas is written only by the composite render pass - nothing else clears it, and its backing texture is allocated uninitialized - but the pass result was ignored and the texture composited regardless, flashing an arbitrary color. The renderer now falls back to the geometric-clip matte path when the composite fails.
  • Lottie: a paint-less nested group now contributes its geometry to the enclosing group's paints with its own modifiers applied. The geometry was rebuilt from raw shapes, dropping the nested group's trim, repeater, merge-paths and rounded-corner modifiers, which is what defines the outline: RubberHose rigs draw a limb as a 4-point star trimmed to a quarter, so the parent stroke painted the whole star instead of an arc (the stray stars in mughead.json and pumped_up.json).
  • Lottie: track mattes (alpha, alpha-inverted, luma, luma-inverted) now composite the matte source's rendered alpha - including its fill opacity, gradients, and anti-aliased edges - instead of hard-clipping the target to the source silhouette. The matte source and target are rendered into offscreen GPU buffers (sized to the fitted on-screen resolution) and multiplied by a fullscreen matte-composite shader. A partially transparent matte source now shows through correctly (e.g. matte_two_item_with_lowerlayer.json, whose 65%-opacity source blends the white matted ellipse to pink over the red layer beneath). Falls back to the previous geometric-clip behaviour when no GPU is available (e.g. headless rendering).
  • Lottie: EllipseShape paths now start at the top (12 o'clock) and follow the shape direction (clockwise for d == 1, counter-clockwise for d == 3), matching Lottie's convention. Previously they started at the right (3 o'clock) going counter-clockwise, which placed trimmed arcs at the wrong position (e.g. the expanding rings in world_locations.json were cut short on the right).
  • Path::withRoundedCorners() left one corner sharp on closed subpaths whose geometry ended with an explicit segment back to the start vertex (as produced by Lottie bezier toPath()). The duplicated start/end point formed a zero-length edge that made that corner degenerate. The trailing duplicate is now dropped, and corners are rounded with a cubic arc (circle kappa) instead of a single quadratic through the vertex, so a square with a full Round Corners modifier becomes a proper circle (e.g. the morphing loader shape in loader.json).
  • Lottie: trailing top-level modifiers (trim, repeater, rounded-corner) in a shape layer now apply to every preceding top-level group in the run, not just the last one, so a single trim animates all shapes it should (e.g. the knife in it's_lunch_time!.json, and the segmented strokes in imprint.json / fingerprint_success.json). Trailing paints similarly reach all preceding paint-less groups.
  • Lottie: animated properties driven by an AfterEffects loopOut('cycle') expression (AnimationProperty<T>::LoopMode) now repeat their keyframe range instead of freezing on the final value once playback passes the last keyframe. Fixes pulsing markers vanishing after their first cycle (e.g. the orange location circles in world_locations.json).
  • Lottie: precomposition layers are now rasterized to an offscreen texture sized to the on-screen device resolution instead of the fixed composition size, so precomps no longer look blurry when the animation is scaled up (e.g. tractor.json).
  • Lottie: layers with partial (animated) opacity are isolated into a transparency layer for correct compositing; this offscreen buffer is now sized to the fitted on-screen resolution instead of the composition size, so small compositions no longer look blurry when scaled up (e.g. spin,_lil_loader_v2.json, a 90x90 composition whose fading "stick" layers were rasterized at 90px and upscaled).
  • Lottie: Merge Paths (mm) is now supported (AnimationMergePaths). Boolean modes (Add/Union, Subtract, Intersect, Exclude) combine the preceding path geometry with the matching boolean operation, while the plain "Merge" mode concatenates paths and lets the fill winding rule form counters (holes). Nested paint-less groups only feed their geometry to the parent group's fills/strokes when a Merge Paths modifier is present; otherwise nested groups stay self-contained so paint-less construction guides are not accidentally filled (fixes stray star/cross shapes and per-frame overhead in pumped_up.json and mughead.json). Fixes shapes built from merged sub-paths rendering only partially (e.g. the red windmill sails in windmill.json) without filling in letter counters (e.g. the holes in "O"/"A" in goal.json).
  • Lottie: the AfterEffects inertial-bounce ("overshoot") position expression (amp/freq/decay) is now approximated via AnimationTransform InertialBounceParams, producing the decaying oscillation past the last position keyframe. Fixes elements that dropped in without the expected bounce (e.g. windmill.json).
  • Lottie / AnimationRenderer::renderComposition: content that extends beyond the composition viewport (e.g. shapes with coordinates outside the w/h bounds, as in jolly_walker.json) now clips to the fitted composition rectangle instead of the full target bounds, so it no longer spills into the letterbox / pillarbox area when the target rectangle is not the composition's aspect ratio.
  • AnimationTransform::positionAt() spatial bezier motion paths were nearly straight instead of curved: the second control point used the next keyframe's incoming tangent (k1.tangentIn) rather than the current segment's own tangent (k0.tangentIn). In Lottie both to and ti belong to the keyframe starting the segment, so a circular motion path (e.g. a shape orbiting on a bezier arc) collapsed toward linear interpolation.
  • OpenGL / WebGL: the main frame's rive flush went silently blank (draws degenerate, screen frozen on the last good frame) whenever a GpuCanvas committed mid-frame. endOffscreen()'s unbindGLInternalResources() wipes the shared GL texture units, but the main render context's internal textures (tessellation/gradient/feather/atlas) were only rebound at begin() - before paint() - so any offscreen 2D flush during paint left the main flush sampling incomplete textures (no GL error; GLES returns zeros). The GL backend now calls invalidateGLState() on the flushing context immediately before every flush() (main frame and offscreen), making each flush self-contained regardless of how many rive/ore contexts interleave on the one real GL context. Fixes SpinningCubeDemo on WASM/WebGL2 appearing frozen (with sporadic 5-15 s updates) and the page turning sluggish while the app still reported ~57 FPS.
  • Graphics::drawTexture / drawImage / transparency layers rendered nothing (transparent) whenever the rive frame ran in atomic interlock mode - always the case on the iOS simulator, and on any platform when raster ordering is disabled. The composite was implemented as a path draw with an image paint, which atomic-mode shaders cannot sample; Graphics::renderTexture now routes through rive::Renderer::drawImage, which falls back to a dedicated image-rect draw in atomic mode. Fixes invisible Lottie precomps/mattes, GpuCanvas composites, and the SpinningCube demo output on the iOS simulator.
  • OpenGL / WebGL: GpuCanvas textures drawn with Graphics::drawTexture (Lottie precomp caches and matte composites) rendered vertically flipped, because the GL canvas source texture is stored bottom-up. GpuTexture::getOrAdoptGpuTexture() now prefers the Y-flipped sampled mirror - kept fresh at each canvas flush - matching what GpuRenderPass already did for sampled inputs. No change on Metal/D3D, where the mirror is null.
  • SDL3 windowing: mouse move/drag was broken on touch platforms (iOS, Android). Motion was synthesized only by polling SDL_GetGlobalMouseState, which has no backend implementation there and falls back to window-relative coordinates, so subtracting the window position shifted every move. Touch platforms now consume the touch-synthesized SDL_EVENT_MOUSE_MOTION events directly; desktop keeps the global-cursor poll (needed for embedded plugin editors).
  • SDL3 windowing: mouse drag events were lost inside embedded plugin editors (notably on macOS, where the host owns the native application so SDL never receives Cocoa mouse focus and suppresses drag motion). Dragging is now synthesized by polling the global cursor while a button is held, on the message thread, for all platforms.
  • Slider could get stuck showing its hover color after a touch drag: mouseEnter/mouseExit never fire for touch (no hover phase, and drag capture bypasses them for the mouse too), so releasing outside the slider's bounds left the hover state on. mouseUp now clears it directly when the pointer isn't over the slider anymore.
  • UBSAN and ASAN fixes throughout the codebase
  • AUv3 plugin host bypass is now connected to the processor: the wrapper-owned bypass parameter is created and drives processBlockBypassed, and host bypass state is persisted/restored inside the YUPProcessorState blob (legacy raw processor state still loads)
  • Added bypass parameter handling tests for the AU, CLAP, and VST3 plugin client wrappers (routing to processBlockBypassed, bypass state round-trip, and text/value conversion)
  • Windows toasts emit the scenario attribute with the spellings the toast schema declares (reminder / alarm / incomingCall) rather than the capitalised WinToast ones, which are not part of the enumeration. Schema conformance only — it is not the cause of the toasts that fail to display on Windows 11, see docs/Windows Toast 80070490 Analysis.md
  • Windows toasts report a real permission state instead of always claiming granted: ToastNotification::getPermissionState() / requestPermission() now query IToastNotifier::get_Setting(), so an application, user, group policy or manifest level block is visible to the caller. The setting is also logged next to the payload. Note that it does not cover Do Not Disturb or the per-app "show notification banners" switch, which suppress the on-screen banner while still delivering the toast to the notification center
  • Windows toasts no longer hand put_ExpirationTime a stack object that dies at the end of the enclosing if block. The notification retains that IReference<DateTime> for its whole life, so it was already dangling by the time Show() read it; it is now a reference-counted ComBaseClassHelper that the notification keeps alive
  • TypeErasedObject now relocates its payload through the payload's move constructor instead of a byte copy. Moving a payload that points into itself - libstdc++'s small-string std::string, a std::map header, or anything caching a member's address - left those pointers aimed at the dead source buffer, so GpuPipeline::Impl's entry-point strings freed a stale stack address when the pipeline was destroyed. glibc reported it as free(): invalid pointer and killed yup_tests in GpuAttachmentMockTests on Linux, while libc++'s std::string has no self-pointer, so macOS never saw it

Documentation

  • Added a dedicated DSP documentation area (docs/dsp/) covering yup_dsp end to end: math/windowing/noise, FFTs and spectral analysis, filter design and filter implementations, dynamics and metering, onset detection, convolution and delay, resampling, and time-stretching/pitch-shifting
  • docs/dsp/yup-dsp-language.md (the YDSP language reference) is now linked from docs/dsp/index.md's toctree - it previously built but was unreachable from the docs site. Its §2.7 EBNF now covers the bitwise operators (& | ^ ~ << >>, at their actual precedence, which is tighter than comparisons unlike C) that were already implemented but undocumented; §2.8 lists the previously-undocumented asinh/acosh/atanh/round/copysign intrinsics and the new integer overload of min/max/clamp/abs/sign (including the abs(INT_MIN) edge case); and §2.7/§3.2 each gain a sentence clarifying that unary ~ (bitwise not) and the graph algebra's binary ~ (recursion) are unrelated operators in separate grammars, not an overload of one operator.

[1.0.0] - 2026-07-03

Platform Support

Android

  • Android window support with YupActivity Java class (#29, #34)
  • Java bytecode compilation via yup_android_java.cmake (#53)
  • External storage permissions (READ_EXTERNAL_STORAGE / WRITE_EXTERNAL_STORAGE) for file access (#61)

iOS

  • iOS CI pipeline with Xcode toolchain (#8)
  • Updated minimum deployment targets: iOS/tvOS 13.0, watchOS 6.0, macOS 11.0 (#72)
  • ARC enabled by default on Apple platforms (#91)
  • iOS Simulator-specific framework groups (iosSimFrameworks / iosSimWeakFrameworks) in module declarations (#48)

macOS

  • macOS message loop reworked: time-sliced event dispatch via CFRunLoopRunInMode targeting ~60 Hz; quit event registered with NSAppleEventManager for proper Apple Event quit handling (#47)
  • NSSupportsSuddenTermination = false added to macOS Info.plist (#47)

Emscripten / WebAssembly

  • Full Emscripten/WASM support including AudioWorklet audio device (#25)
  • WASM threading with exported runtime methods (#61)
  • -msimd128 compile flag and configurable stack size for Emscripten targets (#98)

SDL2

  • SDL2 integration with libpng, libwebp, and rive_decoders (#37)
  • Improved SDL/JUCE Message Manager dispatch loop (#38)
  • SDL symbol namespacing to prevent linker conflicts in Apple platform plugins (#112)

Graphics

  • Reworked rendering backend selection API: YUP_RIVE_USE_D3D, YUP_RIVE_USE_METAL, YUP_RIVE_USE_OPENGL, YUP_RIVE_USE_DAWN (#24)
  • Headless graphics context and no-op Rive factory for offscreen rendering (#32, #52)
  • SVG rendering support (#56, #64)
  • SVG 1.1 spec compliance: blend modes, patterns, polygon/polyline (#100, #118)
  • Path API improvements with comprehensive examples (basic shapes, arcs, curves, transforms, advanced) (#56)
  • createStrokePolygon() with feather effects (#55)
  • Improved color management and gradient editor (#87)
  • Improved CPU and GPU image rendering (#39, #87)
  • Color::brighter() / darker() made const (#32)
  • AffineTransform::inverted() constexpr method (#19)
  • constexpr math utilities: juce_abs(), jmap(), jlimit(), findMinimum() / findMaximum(), nextPowerOfTwo(), and more (#18)
  • ColorGradient::Spread enum: Pad, Repeat, Reflect tiling modes with withSpread() builder (#119)
  • CubicBezier class: pointAt(), derivative(), length(), splitAt(), bounding box, and intersection (#119)

Image Formats

  • New image format I/O framework: ImageFormat, ImageFormatReader, ImageFormatWriter, ImageFormatManager - plugin-style registry with magic-byte detection and animated image support (#119)
  • BMP image format: reader (1/4/8/16/24/32-bpp, RLE4/RLE8, palette) and writer (24-bpp uncompressed), controlled by YUP_IMAGE_FORMAT_BMP (#119)
  • PPM/PGM/PBM (Netpbm) image format: full P1–P6 plain and binary read/write, controlled by YUP_IMAGE_FORMAT_PPM (#119)
  • PNG image format via libpng: grayscale, grayscale+alpha, RGB, and RGBA at 8- and 16-bit depths, controlled by YUP_IMAGE_FORMAT_PNG (#119)
  • JPEG image format via libjpeg: quality-level encoding, controlled by YUP_IMAGE_FORMAT_JPEG (#119)
  • WebP image format via libwebp, controlled by YUP_IMAGE_FORMAT_WEBP (#119)
  • Animated GIF image format via libgif: per-frame delay, loop count, animated write API (beginAnimation / writeFrame / endAnimation), controlled by YUP_IMAGE_FORMAT_GIF (#119)
  • Image::loadFromData() reimplemented via ImageFormatManager (#119)

Offscreen Rendering

  • GraphicsContext::OffscreenTarget abstract interface for opaque platform GPU offscreen resources (#119)
  • Offscreen API on GraphicsContext: createOffscreenTarget(), beginOffscreen(), endOffscreen(), readOffscreenPixels() - implemented for Metal, OpenGL, and D3D backends (#119)
  • Graphics constructors for rendering to an Image or OffscreenTarget outside the main frame cycle (#119)
  • Image gained renderCanvas backing (RenderCanvas) alongside texture for offscreen render-to-texture; duplicate() re-enabled with proper deep copy (#119)
  • Graphics::TransparencyLayer RAII class for isolated group opacity compositing: renders into an offscreen target and composites back at the given opacity on commit() (#119)

Animations (yup_animation)

  • New yup_animation module: Lottie-compatible animation engine depending on yup_core and yup_graphics (#119)
  • Animation: high-level handle with loadFromFile(), loadFromData(), loadFromStream(), renderFrame(), renderAtTime(), renderAtProgress(), toJson(), and saveToFile() (#119)
  • AnimationPlayer: stateful playback controller with forward, reverse, and ping-pong direction modes, looping, variable speed, frame-range clamping, seek, and onFrameChanged / onLoopCompleted / onPlaybackEnded callbacks (#119)
  • AnimationEasing: cubic bezier easing with named presets (linear, easeIn, easeOut, easeInOut, hold) and fromLottieTangents() import (#119)
  • AnimationProperty<T>: generic animated property with keyframe interpolation; specializations for float, Point<float>, Size<float>, and Color (#119)
  • AnimationTransform: animated anchor, position, scale, rotation, and opacity with conversion to AffineTransform at a given frame (#119)
  • Full animation data model: AnimationComposition, AnimationGroup, AnimationLayer, ShapeLayer, shape types (ellipse, rect, path, star, merge, trim, repeater, polystar), paint types (fill, stroke, linear/radial gradient), and modifiers (#119)
  • LottieReader: parse .json and .lottie (ZIP) files from file, string, or stream into the animation data model (#119)
  • LottieWriter: serialize the animation data model back to Lottie JSON (pretty or compact) with full round-trip support (#119)
  • LottieExpressionEvaluator: JavaScript expression evaluator for Lottie property expressions via JavascriptEngine (#119)
  • AnimationRenderer: renders an AnimationComposition to a Graphics context - layer hierarchy, parent-child transforms, matte layers (track-matte), shape fills/strokes/gradients, and image layers (#119)
  • AnimationFrameExporter: exports individual or all frames to Image objects via offscreen GPU, and exports animations to animated GIF files (#119)
  • AnimationRenderer no longer renders a precomposition into an offscreen GPU target unless more than one layer shares that asset: a single-reference precomp now draws straight into its parent, removing a render target, its clear and a full GPU flush per nesting level. Nested precomps previously each opened their own target
  • Fixed AnimationRenderer rebuilding a precomposition's viewport clip with a path boolean op on every layer: the renderer already intersects two rectangular clips itself (in float space, through its clip-rect path), and the boolean op turned the viewport into a polygon with a redundant vertex per crossing that then had to be re-tessellated as a clip path. Clips that cannot overlap the one in effect now cull the layer or precomp outright
  • Fixed the per-layer mask clip cache never being reusable while playing back: it was keyed by frame number even for masks that never animate, so a static mask re-ran its boolean ops on every frame. A static mask is now cached on the layer itself (keyed by composition size, which the mask bounds derive from), and an animated one still caches per frame
  • AnimationRenderer's parent-transform resolution is now linear: every layer's accumulated transform is resolved once and memoized, instead of rescanning the whole layer list (and re-testing every unresolved parent) until nothing more resolves. A cyclic parent chain is bounded rather than retried
  • AnimationRenderer no longer isolates a layer behind an offscreen composite when that cannot change the result: a solid, image, text or null layer at partial opacity already folds its opacity into its single paint, so compositing it through a render target reproduced the same pixels at the cost of a full render target, a clear and a GPU flush per layer per frame. Layer types that draw several overlapping primitives, and any layer carrying a drop shadow, still composite offscreen
  • Fixed AnimationRenderer isolating a layer over a full-composition offscreen target: the target is now sized to the layer's own content box (intersected with the composition viewport) and the content shifted to match, so a small layer in a large composition no longer pays for the whole composition in clear, flush upload and composite. Rasterization resolution is unchanged
  • Opacities of 0.999 and above are now treated as fully opaque. Exporters write values like 99 or 99.9 for layers meant to be seen at full strength, and isolating one behind an offscreen composite to reproduce a sub-1% difference in alpha cost more than the difference was worth
  • Fixed AnimationRenderer treating a precomp as shared when its referencing layers are never on screen together: reference counting now counts only the layers the frame actually draws (hidden, matte-source and out-of-range layers excluded, each nested level evaluated at its own frame). A composition split into sequential time slices - a common export shape - is therefore drawn like any single-reference asset instead of through a full-size offscreen target on every frame

Audio

Plugin Support

  • VST3 plugin support (#44)
  • Audio Unit (AUv2) plugin support (#93, #106)
  • Barebone Audio Unit (AUv3) plugin support (#122)
  • Barebone AAX plugin support (#122)
  • Barebone LV2 plugin support (#122)
  • Audio Plugin Host for AUv2, VST3, and CLAP (#93, #98, #106)
  • Standalone plugin support with improved audio parameters (#46)
  • CLAP/VST3/AU validators and code signing (YUP_ENABLE_VST3_VALIDATOR, etc.) (#106)
  • pluginval integration for automated VST3 validation (#67)
  • Sidechain and multi-bus audio input support across VST3, CLAP, AUv3, AUv2, AAX, and LV2: AudioBus gains a Role (Main/Auxiliary) and isDefaultActive, AudioProcessContext exposes per-bus inputs/outputs views (AudioBusBufferView) with getMainInput()/getAuxiliaryInput()/getMainOutput() accessors, and secondary input buses are forwarded to the processor instead of being discarded

Audio Formats (yup_audio_formats)

  • New yup_audio_formats module: AudioFormat, AudioFormatManager, AudioFormatReader, AudioFormatWriter, WAV codec (#51)
  • Opus, MP3, FLAC, AAC, CoreAudio, and WMF codec support (#86, #88)

DSP (yup_dsp)

  • New yup_dsp module with FFT/windowing via ooura, pffft, vDSP, IPP, FFTW3 (#71)
  • Basic IIR filter implementations (#71)
  • Linkwitz-Riley crossover filters (#71)
  • FIR filter (#75)
  • Partitioned convolution (#75)
  • Oversampler and Resampler (2×/4×/8×) (#97)
  • Noise generators (#71)
  • Virtual analog filters: AnalogTwoPoleFilter, AnalogVowelFilter, AnalogKorg35Filter, AnalogMoogLadderFilter, AnalogRolandDiodeFilter, CombFilter (#103)
  • Spectral processor (#116)
  • Onset detectors (SpectralFlux and ComplexFluxODF) with perceptual filter bank (#117)
  • Time-domain and Frequency-domain stretching with backend selection (homebrew PSOLA plus bungee) (#104)
  • Distortion processors with oversampling: TanhDistortionProcessor, BlunterSoftClipperProcessor, AaIirHardClipperProcessor (#108)
  • Click-less fractionally addressed delay (FAD) (#109)
  • Emscripten AudioWorklet audio device (#25)
  • MIDI 2.0 / Universal MIDI Packets (UMP) implementation (#83)
  • KMeterState de-interleaving buffers pre-allocated in prepare(), eliminating per-block heap allocation in the audio callback (#119)

Synthesiser

  • SynthesiserVoice converted to ReferenceCountedObject with Ptr = ReferenceCountedObjectPtr<SynthesiserVoice>; Synthesiser::addVoice() now accepts SynthesiserVoice::Ptr (#82)

Audio Graph (yup_audio_graph / yup_audio_plugin_host)

  • New yup_audio_graph and yup_audio_plugin_host modules (#93)
  • Thread-safe BufferingAudioSource with atomic nextPlayPos (#98)
  • AudioPlayHead::getContinuousTimeInSamples(): continuous sample time without loop-reset (#35)

UI

Components and Widgets

  • Customizable theming/skinning system (#13)
  • mouseDoubleClick() virtual method with configurable threshold (#14)
  • KeyboardFocusMode enum replacing boolean setWantsKeyboardFocus(); added textInput() callback (#16)
  • PopupMenu and ComboBox components (#57, #62)
  • Native file chooser via FileChooser (#61)
  • Component paint profiling: PaintProfiler with ring-buffer stats (min/max/mean/p50/p95/p99) (#95)
  • ComponentNative::getGraphicsContext() virtual method allowing components to access the GPU context for offscreen operations (#119)
  • MouseListener weak-referenceable interface for all mouse events; Component::addMouseListener() / removeMouseListener() (#30)
  • Improved slider components (knob, linear, range) and button components (#70)
  • Unified drag-and-drop support in Component: isInterestedInDrag() / itemsDropped() virtuals with a fluent DragAndDropData payload (files and text on SDL, URIs reserved for future backends); drops dispatch to the topmost interested component and bubble up to parents
  • Added unit coverage for SystemClipboard data formats and Component drag-and-drop callbacks
  • Safe area support: Component::getSafeAreaBounds() and safeAreaChanged() virtual (backed by ComponentNative::getSafeAreaBounds()), so content can avoid display cutouts and system bars on mobile devices
  • High dpi support on Windows and Linux X11: window bounds, screen geometry and input coordinates are now logical points everywhere (converted at the SDL boundary), so windows and content scale with the display scale like on macOS; live display scale changes resize the native window keeping the logical size

Text

  • TextEditor and Label components (#16, #55)
  • Improved fonts: better layouting, variable font axis manipulation, embedded fallback font (#55)
  • Clipboard support: text, MIME-typed data with lazy callbacks, and primary selection (#55)

Audio GUI (yup_audio_gui)

  • New yup_audio_gui module (#70)
  • MIDI keyboard component (#70)
  • Filter frequency response visualisation (#71)
  • Spectrum analyser component (#71)
  • Spectrogram component with peak/RMS/power/PSD level modes (#102)
  • Oscilloscope and spectrum analyzer display processors (#109)

Data Models (yup_data_model)

  • New yup_data_model module (#15)
  • UndoManager with Transaction, ScopedTransaction, UndoableAction (#15)
  • DataTree hierarchical data structure with builder pattern and transactional mutations (#73, #74)
  • DataTree query support (#74)
  • DataTree schema validation (#74)
  • CachedValue<T> for type-safe DataTree property references (#74)
  • DataTree complete UndoableAction suite for transactional mutations: PropertySetAction, PropertyRemoveAction, RemoveAllPropertiesAction, AddChildAction, RemoveChildAction, RemoveAllChildrenAction, MoveChildAction, CompoundAction - all with full undo/redo semantics (#82)
  • Identifier usable as std::unordered_map key via std::hash specialization (#27)

Artboard (Rive Integration)

  • Improved artboard placement (#17, #43)
  • Shared Rive file across multiple Artboard components (#43)
  • State machine inputs: setNumberInput(), setBoolInput(), triggerInput() (#17, #43)
  • advanceAndApply() and durationSeconds() for timeline control (#17)
  • State machine event handling (#43)
  • Multi-artboard component support (#43)

Core & Utilities

New Modules

  • yup_simd: SIMD vectorization framework (#107)
  • yup_python: Python bindings (from popsicle) (#65)

yup_core Additions

  • constructAt() / destroyAt() / voidify() in memory/yup_Memory.h: portable replacements for std::construct_at / std::destroy_at, used by TypeErasedObject
  • TypeErasedObject now supports class template argument deduction (deduction guide sizes storage to the stored value) and move construction / assignment from a smaller-sized TypeErasedObject
  • SqliteDatabase with Statement and Transaction (#94)
  • Perfetto profiling: YUP_ENABLE_PROFILING, Profiler singleton, YUP_PROFILE_START / YUP_PROFILE_STOP macros (#20)
  • Watchdog file watching utility (#50)
  • URL copy and move constructors (#58)
  • messageThreadID made atomic (removed mutex) (#26)
  • ReferenceCountedObject::incReferenceCount() / decReferenceCount() made const (#28)
  • DatagramSocket multicast overloads with local IP (#60)
  • ResultValue<T>::valueOr() (#96)
  • AudioSampleBuffer::fill() overloads (#110)
  • JavascriptEngine::executeWithResult() to execute a code block and capture the last expression result (#119)
  • JavascriptEngine::registerNativeFunction() for top-level native function registration by name (#119)
  • JavaScript $ accepted as valid identifier character, required for Lottie expression compatibility (#119)
  • WASM: POSIX file API extended with symlink(), dirent.h, fnmatch.h, utime.h support; WASMFS enabled for standalone builds (#36, #59)
  • Linux: File::isOnRemovableDrive() implemented via /sys/block/<dev>/removable (#36)

Build System

  • zlib and oboe extracted from inline module sources into standalone thirdparty/ modules for cleaner namespace isolation and build separation (#6, #7)
  • TARGET_IDE_GROUP parameter on yup_standalone_app and yup_audio_plugin; all modules, tests, and examples placed in dedicated "Modules", "Tests", and "Examples" IDE folders (#11)
  • appleFrameworks / appleWeakFrameworks module declaration fields unifying iOS and macOS framework lists (#36)
  • Per-platform C++ standard override via *CppStandard module header fields (appleCppStandard, osxCppStandard, linuxCppStandard, wasmCppStandard, androidCppStandard, msftCppStandard) (#52)
  • Platform CMake files reorganized under cmake/platforms/ and loaded dynamically per target platform (#36)
  • Circular dependency detection for YUP modules (#111)
  • Module link options support (per-platform *LinkOptions) (#53)
  • Module target aliases (yup::yup_core, etc.) (#53)
  • Code coverage: YUP_ENABLE_COVERAGE, codecov integration (#54)
  • Test sharding support: --gtest_total_shards / --gtest_shard_index for parallel CI runs (#119)

Bug Fixes

  • Crash at startup when height/width is 0 on custom-scaled screens (#21)
  • Application never quits: incorrect quitMessagePosted ordering in stopDispatchLoop() (#42)
  • Redraw issues and app icon rendering on macOS (#31)
  • iOS toolchain: removed hardcoded DEVELOPER_DIR path (#72)
  • CoreAudio thread safety: atomic operations replacing mutex (#76)
  • SMPTE timecode validation, SSE macro, and memory fixes (#78)
  • ZIP timestamp: missing >>1 for 2-second resolution (#92)
  • StyledText::clear() fully resets state; caret bounds and glyph index for empty lines (#96)
  • Duplicated SDL symbols in Apple plugins (#112)
  • ComboBox popup re-opens on click-to-dismiss; added ignoreMouseDownAfterPopupDismissal (#114)
  • AudioDeviceManager destructor race on midiCallbackLock (#115)
  • Android oboe: __ANDROID__ preprocessor instead of ANDROID (#9)
  • Graphics::drawImage(), renderStrokePath(), renderFillPath(), renderFittedText(): opacity not propagated to renderer - fixed (#119)
  • SIMDRegister: tail-loop bounds check preventing out-of-bounds access in load and store paths (#119)
  • Mouse-wheel events now dispatched to the component under the cursor when no component is focused (#30)
  • Linux Watchdog: inotify fd set to non-blocking, read buffer heap-allocated, thread join order corrected to prevent crash on destruction (#36)