All notable changes to the YUP project will be documented in this file.
The format is based on Keep a Changelog.
-
MidiKeyboardComponent: multitouch (each finger holds, slides and releases its own note),[<]/[>]octave scroll buttons replacing the-/+selector, a visible range (setVisibleRange,setHighestVisibleKey,onVisibleRangeChanged) inside the available range, and the JUCE velocity sensitivity, channel mask, orientation, black note proportions, key hooks and note label APIs. Behavior change: typed notes follow the newsetKeyPressBaseOctave(setOctaveForMiddleConly names octaves), wheel scroll and zoom stay inside the available range, position velocity is scaled bysetVelocityand correct for vertical keyboards, a default keyboard shows the scroll buttons,focusLostkeeps mouse and touch notes held, vertical keyboards anchor black keys at the back edge, and the back-edge shadow and front-edge line now paint over the white keys instead of under them. -
ComponentNative::setVsyncEnabled()/isVsyncEnabled()switch vsync at runtime (GL swap interval, falling back from adaptive to plain vsync when the driver refuses it, Metal, D3D; Dawn keeps its creation-time mode). On desktop, vsync on runs at the display rate and vsync off paces to the desired frame rate. On the web, vsync on renders on every display refresh and vsync off paces to the desired frame rate on elapsed time, so 60fps on a 144Hz display averages 60 instead of snapping to 72. The Emscripten page shell gains a VSync toggle and an FPS field (greyed and showing the measured rate while vsync is on) -
Drawable: a<use>without its own fill or stroke now draws the referenced shape with that shape's gradient, pattern, opacity, filter, clip and markers instead of a flat copy of its color.<use>of<symbol>or<defs>content now renders. On the rive MSAA backend (WebGL2), a draw faded withGraphics::setOpacityis no longer treated as opaque, so solid colors and opaque gradients blend instead of replacing what is below (any component, not only SVG). -
GpuComputePipeline::compileFromBundleresolves themain0kernel SPIRV-Cross emits for a GLSLmainon Metal, likeGpuPipeline::compileFromBundlealready did, instead of failing with "Metal compute function not found: main". -
CMake:
yup_add_shader_bundleaccepts aCOMPUTEstage, and withBUNDLE_RESOURCEships the.yslthroughBUNDLE_RESOURCESinstead of embedding it. Bundles are only regenerated when the tool, the arguments or an input (stages and the newDEPENDS) changed. The graphics example declares the modules, data and precompiled shaders of each demo, so a single-demo build (YUP_EXAMPLE_GRAPHICS_DEMO) only links what that demo uses. Only the SpinningCube and GpuAudio demos, which edit shaders live, still need the shader transpiler. -
Behavior change
SystemStats:isOperatingSystem64Bitreports the OS instead of the build (32-bit builds on 64-bit Linux, Android, Windows on ARM; iOS now true; best-effort host probe on WebAssembly). AddedMacOS_15,MacOS_26andMacOS_27. The macOS name readsmacOS <version>, the Android version isBuild.VERSION.RELEASEand the device description drops the serial, the Linux version is the kernel release, andWebBrowserno longer aliases theMacOSXbit (compare it with==, or test the family with& WASM). -
Behavior change
Graphics::setClipPath: the clip is now in local coordinates like every draw call, including the drawing area offset, andgetClipPathreturns it in the current local space. Code that mapped the clip bygetTransform().translated (getDrawingArea().getTopLeft())or by top-level bounds must drop that compensation. Clips on components with an offset parent (for exampleDrawable<image>clips) now land where they are drawn. -
Component: a rotated or sheared component, and everything inside it, is now clipped to its real outline instead of the bounding box of its transformed bounds, so children no longer spill past the corners of a skewed parent. -
SyncSpectralResampler: the per-harmonic accumulation goes throughFloatVectorOperationsagain instead of a hand-writtenSIMDRegisterloop, which was many times slower in debug builds. The graphics synthesizer example caps a per-voice synced series at the note's Nyquist harmonic count. -
Graphics synthesizer example: PRISM logo and larger buttons in the header, LFO / scope columns aligned with filter / envelopes, and pitch bend and mod wheels beside the keyboard. The mod wheel is a new
MOD WHEELsource in the modulation matrix. The matrix now has 16 slots. -
CMake: retry failed upstream module and validation tool downloads, verify
sha256after download, and fail at configure time when an upstream archive extracts nothing instead of later with missing headers. -
YDSP backend (emscripten): mint kernel handles from a module-wide counter instead of a JS-realm-local one, so a realm that runs a graph can no longer find another realm's kernel under the same key and silently invoke the wrong module.
-
YDSP VS Code extension: audition the active patch through
yup_dsp_compiler runfrom a Patch Player sidebar view (transport, workspace patch list, audio/MIDI device selects, sample rate, block size, test note), with a single pinned player per window, a status bar, a dedicated playback output channel and an opt-in follow-active-patch mode. -
YDSP player:
--hotreloadnow works for a standalone.ydspas well as a.ydsp-project, driven byYdspDiagnostics::getSourceIds(), which reports the source closure of the last compile (root source, project sources and transitively imported files). -
YDSP diagnostics: check unused library and processor function bodies during editor validation, including return expressions.
-
YDSP tooling: validate standalone processor/function libraries without a graph; add player error details, verbose device/activity reporting, and a test-note option for diagnosing silent playback.
-
YDSP tooling: compile project bundles, validate unsaved project imports in VS Code, and add audio/MIDI device selection and project hot reload to the command-line player.
-
YDSP: add YAML .ydsp-project manifests with patch metadata, explicit-import source inventories, and overridable processor or graph entry points.
-
YDSP: keep the source value of a float literal that adapts to a
float64context, instead of rounding the constant throughfloat32first; regenerate bundles for codegen revision 14. -
YDSP: preserve qualified function calls when nesting library imports, including calls from processors and other library functions.
-
YDSP: preserve source ranges through diagnostics and imports, print path:line:column with five-line context and caret underlines, and improve malformed-number and parser errors.
-
YDSP: add integer/boolean
matchstatements with scoped arms, single selector evaluation, optional_fallback, constant-arm elimination, and invalid-pattern diagnostics; regenerate bundles for language version 4/codegen revision 13. -
YDSP: add opt-in
trace("value={x}")with bounded allocation-free recording, off-thread formatting/console printing, and native/WebAssembly bundle support; regenerate bundles for ABI 2/codegen revision 12. -
YDSP examples: restore wrapping noise sequences in DigitalDrums, WaveLab and ControlRateWah after the integer saturation change.
-
YDSP: saturate integer kernel add/subtract/multiply across native, WebAssembly and constant folding while retaining direct arithmetic for proven-safe operations; regenerate bundles for codegen revision 11.
-
YDSP: eliminate array bounds checks for proven nonoverflowing integer products and strided indices; regenerate bundles for codegen revision 10.
-
YDSP: preserve integer storage width during constant folding so arithmetic and subsequent comparisons agree with runtime execution; regenerate bundles for codegen revision 9.
-
YDSP: saturate float-to-integer conversions and map NaN to zero across native/WebAssembly and constant folding, preserving direct instructions for proven-safe casts; regenerate bundles for codegen revision 8.
-
YDSP: respect source and destination widths in constant-folded numeric conversions and leave exceptional float-to-int casts unfolded; regenerate bundles for codegen revision 7.
-
YDSP: make constant-folded integer shifts match native/WebAssembly operand widths and masked counts without adding runtime checks; regenerate bundles for codegen revision 6.
-
YDSP: remove proven ring-counter and wrapped delay-tap bounds checks after auditing initialization and event writes; regenerate bundles for codegen revision 5.
-
YDSP: remove redundant bounds checks proven by scalar integer masks, clamps, min/max, selects and nonoverflowing arithmetic; retain checks for mutable or unknown ranges.
-
YDSP: fuse shared comparisons into native selects when every consumer preserves the operands, avoiding temporary bounds booleans.
-
YDSP: use safe-index selection for branchless checked reads from nonempty fixed state arrays.
-
YDSP: compact array bounds checks to unsigned comparisons and hoist delay clamps preceding guarded accesses.
-
YDSP: allow unconditional invariant hoisting and unrelated branch optimization in guarded kernels, preserving short-circuit and bounds checks; enable benchmark ASM dumps with
YUP_DUMP_KERNEL=1. -
YDSP: fold constant int32-to-int64 sign extension so widened constants participate in integer arithmetic folding.
-
YDSP: saturate integer kernel negation, absolute value and overflowing division at both widths in native/WebAssembly code and constant folding; regenerate bundles for codegen revision 4.
-
YDSP: preserve exact signed 64-bit literals, reject overflowing structural sizes and explicit nonfinite constants; regenerate bundles for codegen revision 3.
-
YDSP: preserve vectorization for proven in-range
blockSize - Nstream loops; retain guards when fixed loop and buffer lengths can differ. -
YDSP: guard dynamic state-array, struct-field and indexed-stream accesses, diagnose constant invalid indices, and reject overflowing declared state layouts before flattening.
-
YDSP: lower logical operators and ternaries with short-circuit branches, preserving eager
select()and reporting conservative optimization restrictions. -
YDSP: replace
input value/output valuewithinput parameter/output parameterin processors and graphs, freeingvalueas an identifier, with migration diagnostics and updated examples/editor support; regenerate language-version-2 bundles. -
YDSP: fix kernel listing histograms misreading hex-like mnemonics as machine-code bytes; exclude labels and assembler directives.
-
YDSP: retain cached state-array loads across provably disjoint stores using the same index, with type, overlap, and index-redefinition coverage.
-
YDSP: strengthen scalar tanh accuracy and feedback regression coverage; retain libm after rejecting a bounded ARM64 approximation that slowed the tanh shaper.
-
YDSP: make array store-to-load forwarding type- and lane-aware, recognize disjoint constant-index ranges, and run this cleanup after vectorization/unrolling with overlap regression coverage.
-
YDSP: allow bounded two-iteration unrolling around four-lane native math calls with a conservative call-liveness budget; add output/state parity and rejection coverage.
-
YDSP: expand default-policy timing to all 19 patches and four block sizes, separating prepared processing from reset/setup; add median/spread reporting, run metadata, and previous-JIT log comparison.
-
YDSP: replace the incomplete Zita example with the full steady-state stereo network and EQs; add a local JIT benchmark with reference parameter settings, native listings, impulse dumps, and impulse-tail regression coverage.
-
YDSP: hoist and share vector math call targets across native kernels while preserving strict/fastMath accuracy selection.
-
YDSP: omit unused ARM64 array-base setup and reuse dead local vector FMA addends on ARM64/x64.
-
YDSP: preserve unique temporaries and specialize indices in unrolled loops, enabling subsequent FMA contraction and constant vector array offsets on ARM64/x64.
-
YDSP: reuse state-array reads between writes, encode constant scalar array offsets on ARM64/x64, combine constant multiplication chains under fastMath, and rematerialize dry/wet coefficients after math calls.
-
YDSP: vectorize if-converted stream loops and positive constant starts, preserve fixed-index array dependencies, and hoist iteration-local invariant clamp/conversion chains.
-
YDSP: extend fastMath reciprocal multiplication to shared constant divisors and float64, retaining divisions when the reciprocal is non-normal.
-
YDSP: specialize finite positive constant-base powers as scaled exponentials under fastMath, preserving general power calls in strict mode.
-
YDSP: rematerialize constants used between math calls, reducing call-crossing register pressure in kernels such as the compressor.
-
YDSP: share identical immutable entry constants and coefficients across blocks after loop hoisting, reducing duplicate live values on both native targets.
-
YDSP: use encodable integer immediates for fused comparisons on ARM64 and x64, reducing constant-register pressure in delay-bank kernels.
-
YDSP: fold fused-subtraction write-backs in the shared optimizer and preserve overlapping multiply operands in x64 lowering.
-
YDSP: preserve local initializer snapshots when their source state or local is subsequently assigned.
-
YDSP: fold temporary state-copy chains in the shared optimizer, with coverage for saved values and sample history across blocks.
-
YDSP: share operand lowering between ARM64 and x64, including x64 fused product subtraction and width-correct integer immediates; add emission coverage for both architectures.
-
YDSP: contract product-minus-addend expressions and fold them to ARM64
fnmsub; encode small integer offsets and low-bit masks as ARM64 immediates. -
YDSP: eliminate unconditional loop backedges through comparison-only headers in native codegen; add empty-block coverage and small-kernel block-size benchmarks.
-
YDSP: share compatible input delay taps in one masked history ring, prefer FMA contraction on longer feedback paths, and contract the right-hand product of product differences.
-
YDSP: use indexed optimizer lookups, worklist dead-code elimination and early cleanup convergence; reuse eligible stream scratch by lifetime, simplify delay wrapping, and preserve strict subtraction by negative zero. Add scratch-size reporting, opt-in allocation assertions and compile/dense-event benchmarks.
-
YDSP: retain only
process(const YdspProcessRequest&); migrate examples, tests and benchmarks from positional overloads. -
YDSP: validate processing requests before mutation, add span-based
YdspProcessRequest, return preparation validation failures viaResult, and explicitly reject sample-accurate automation for rate-converted nodes. Slot APIs useBySlotnames to avoid ambiguity with YUP strings. -
YDSP: parameter/meter getters now read atomic block snapshots, with slot-based access and bounded single-consumer parameter draining. Delayed events survive multiple variable-sized blocks; exact-boundary events belong to the next block. Event dispatch uses ordered cursors while preserving equal-offset precedence.
-
YDSP: emit ARM64 and x64 kernels independently of the compiler host, including WebAssembly; version 2 bundles store native/WebAssembly artifacts and their source closure, resolve helpers at load time, and instantiate without disk imports or machine-code regeneration. Version 1 bundles must be regenerated.
-
YDSP:
fmsubFlowers to a single fusedfmsubon AArch64 instead of anfmul/fsubpair.FMSUB d, n, m, acomputesa - n * mwith one rounding, so this removes an instruction, a virtual register and a link from the loop-carried dependency chain of every contractedc - a * b(the ladder filter'sin - fb * z4, the wave shaper's1.0 - env * 0.5). It also settles an inconsistency:lowerFusedMultiplyAddalready gives a target without the instruction one rounding through its float64 expansion, so the FMA-capable target was the less accurate one. Pinned by a numeric test rather than left implicit. -
Tests:
yup_YdspBenchmarkTestsgains nine real-life effect benchmarks - a fractionally addressed feedback echo, an LFO chorus, a Freeverb eight-comb/four-allpass reverb, a dB-domain bus compressor, a tanh drive distortion with tone control, a TPT state-variable low-pass, a six-stage phaser, a Karplus-Strong pluck and a 12:1 sample-hold/bitcrush lo-fi processor. Each compares the JIT kernel against a hand-written C++ routine over the standard benchmark length, across the four optimisation policies, and under the same loose parity guard the other shapes use (checksum where the loop stays libm-free, relative magnitude where it does not). The five effect shapes mirror the shippedfx/example processors with their UI annotations stripped. -
YDSP: loop-invariant code motion now places a hoisted instruction as early in the sample loop's preheader (the kernel entry block) as its own operands allow - right after the last preheader instruction that defines one of them, never before. An invariant that reads no scalar state (a hoisted
tan/expcoefficient) therefore lands ahead of the entry block's scalar-state loads and is no longer crossed by the loop's register-promoted state, which the register allocator previously parked on the stack for the whole block loop; the TPT state-variable filter benchmark dropped from ~1.9x to ~1.25x of its hand-written C++ reference as a result. An invariant that consumes a state register (a voice body'senv * gain) still lands after the load that defines it. -
YDSP: the native backends fuse a comparison into the branch that consumes it. A
branchIfwhose condition is a single-use integer or float comparison now re-emits that comparison into the condition flags at the terminator and branches on them (jccon x86-64,b.condon AArch64), instead of materialising a 0/1 register withsetcc/csetand testing it - every sample-loop header and data branch drops the register round trip. -
YDSP: values live across a call now survive it in callee-saved registers instead of round-tripping through the stack at every call. The bundled AsmJit (thirdparty/asmjit_library) previously spilled unconditionally any value sitting in a register a call clobbers whenever that value's live range crossed a basic-block boundary - which meant a per-sample libm call (sin/tanh/log/pow in an effect chain) stored and reloaded the loop's state, stream pointers and coefficients on every sample.
bin_packnow detects call-crossing values by live-span containment over the call sites and gives them preserved homes up front, packing them ahead of any call-free value so none can claim the last preserved register as a fallback first (x19-x28/d8-d15on AArch64), and the local allocator only parks a value in a preserved register at a call when one is free, spilling only on overflow. Against hand-written C++ references the chorus benchmark dropped from ~2.4x to ~1.1x, distortion from ~1.6x to ~1.12x, tanh shaper to ~0.97x, and the phaser now runs at ~0.68x. The changes are small local patches over upstream AsmJit, marked at the sites incore/ra_pass.cppandcore/ra_local.cpp. -
YDSP: scalar libm function addresses are materialized once per kernel instead of at every call site. A per-sample
sin/tanh/powused to rebuild its 64-bit target register (movz+ twomovks) before each register-indirect branch; the codegen now pre-scans the IR, emits one address load per distinct function at kernel entry (materializeScalarLibmTargets) and lets every call site branch through that shared register, which the allocator keeps in a callee-saved slot. -
YDSP:
foldStateWriteBacksnow folds integer state write-backs too, not just float chains. A per-samplewp = wp + 1ring-pointer update used to be a freshaddIplus amovIwrite-back into the loop-carried register every sample - exactly the floatfmovround-trip the pass already removed - and so did the sample loop's own induction move. Integer arithmetic, bitwise andselectproducers now write the carried register in place (wp = wp + 1is oneadd), dropping one instruction per sample off the state chain in the delay-line, chorus, karplus and lo-fi shapes. The canonical write-back move of a constant-bound loop's induction is preserved (fusion, unrolling and the vectoriser pattern-match the loop body on it before the fold runs); the fold is what removes it after unrolling. -
YDSP: the vectoriser now accepts
blockSize - k(andblockSize + k) loop bounds instead of rejecting them asunsupportedLoopBound. They are runtime bounds like plainblockSize, so the same epilogue machinery widens them tobound & ~(lanes - 1)whole vectors plus a scalar remainder loop, with the stream-access requirement enforced as before; afor i in 0..blockSize - 1stream loop widens the same way0..blockSizedoes. -
YDSP: the vectoriser now widens loops containing the rounding intrinsics
floor,ceilandrint. Each lowers to one native packed instruction (frintm/frintp/frintnon AArch64,roundpson x86), so a block-mode quantiser or bitcrush loop over streams vectorises like any other element-wise loop and stays bit-exact per element. Vectorroundandcopysignremain scalar: round-half-away has no one-instruction packed form on x86 and copysign needs float bitwise sign-mask work both backends do not expose yet. -
YDSP: if-conversion now speculates an input-stream load when it reads the induction of the stream-length loop that directly encloses it (the sample loop, or a block-mode
for i in 0..blockSizeloop). A periodic sample-and-hold or downsample branch (if (counter == 0) { held = in; }) becomes straight-line code that loads unconditionally and selects, instead of a branch that mispredicts every Nth sample - the residual gap in the lo-fi 12:1 sample-hold benchmark. State-array and parameter loads stay inside their branch, as do loads under a constant-bound inner loop. -
YDSP: scalar leaf values that would otherwise be parked around a per-sample libm call are now rematerialized instead. A parameter load or compile-time constant defined before the loop but used only after the loop body's last libm call gets re-defined right after that call (
rematerializePostCallLeaves): it reloads once per sample - the same value - and its live range never spans the call, so the register allocator stops emitting the per-samplestr/ldr [sp]parking pair and one less value competes for a preserved register. Leaf defs only; expF is not treated as a boundary because fastMath inlines it on AArch64. -
YDSP: voice banks now work on subgraphs -
node v = VoiceChain[8]inlines the whole chain as one runtime voice group, so an effect (filter, delay, reverb, shaper) runs once per voice with its own state instead of once after the summed bank. A banked subgraph must declare aninput eventand exactly one float32 output stream and contain at least one event-handler member; members must be ordinary single-voice float32 processors (nested banks, per-member rate changes and midi-only members are rejected). Group members run voice-major with per-voice intra-group delay rings, each note event is allocated once per group and fanned to every subscribing member on the same voice, all-sound-off clears the group's slots once and silences every member, andgetActiveVoiceCountaccepts the bank name.examples/graphics/data/synths/PerVoiceEcho.ydspis the worked example; the crosstalk test proves two notes through the chain are numerically different from the same effect placed after the mix. -
YDSP: a
statearray may now omit its size and let a{ ... }initialiser list determine it -state float wavetable[] = { ... }- and the compile-time pseudo functionsize (...)returns the element count of any array expression: astatearray (explicit or inferred; an array of struct instances yields its instance count), a struct array field (size (comb.buf)/size (combs[i].buf)), and, in a block-mode processor, a stream (whose length is the runtimeblockSize). A[]state without a non-empty{ ... }list, and struct-array states written[], are compile errors with a message naming the fix. -
YDSP: one
statestatement may now declare several states of the same type, separated by commas -state float x, y, z;. Each declarator keeps its own array size, initialiser and trailing annotation (state int active [[ role: voiceActivity ]], released;marks onlyactive), and the list is sugar for the equivalent run of single-statestatements. -
Examples:
fx.Delayis now a fractionally addressable delay. The read tap is a linear interpolation between the two ring samples either side of the requested delay, and itstimeparameter is smoothed, so changing the delay time while audio is running glides the tap instead of clicking on whole-sample steps.PerVoiceEcho.ydspruns one such delay per voice. -
YDSP: adjacent members of one banked voice chain that are pure per-sample processors now fuse into a single kernel, exactly as plain chains always did. The fused member still runs once per voice under the group's slot table (per-voice state and per-voice parameters are preserved), but the intra-group junction stops being a per-voice scratch round-trip; fusion never crosses a group boundary.
-
YDSP: a banked voice chain now skips a whole voice when every member reports its
[[ role: voiceActivity ]]flag asleep - the flag is no longer restricted to processors that declare event handlers, so pure effects (filters, delays, reverbs) can opt in by keeping their flag set until their own tail has died. Skipping freezes the whole voice (per-voice delay rings included); a held voice, a pending event, or block-wide automation/all-sound-off always keep the voice running. Fusion and skipping are complementary: fused effect members cannot carry a flag, so their voices always run. -
YDSP: the optimizer now performs block-local common-subexpression elimination for pure IR expressions, reducing repeated arithmetic without changing non-SSA state or memory semantics.
-
YDSP: endpoint annotations gain
[[ mid: <value> ]](the value that should sit at the middle of a host slider's travel, exposed asYdspParameterInfo::midValueso a UI can derive a logarithmic skew) and[[ bipolar: true ]](a range-centered-on-zero flag, exposed asYdspParameterInfo::bipolar, defaultfalse).examples/graphics's YDSP Synth Lab appliesmidthroughSlider::setSkewFactorFromMidpoint(); the Analog Saw patch's Cutoff knob demonstrates it. -
YDSP: introduced the
.ydsbbundle API,yup_dsp_compilerhost tool, CMake embedding helper, and bundle format documentation. -
YDSP: fast-math contraction now covers multiply-add and subtract-multiply patterns; the explicit tradeoff is changed rounding. Fixed inline
@delays use a compact increment-and-wrap IR operation for faster native code. -
YDSP:
YdspCompilernow accepts per-compileYdspCompileOptions, providing baseline/automatic/aggressive policies, strict-by-defaultfastMath, host or portable-target selection and an optional optimisation report. Native bank-loop SIMD now uses the selected target width (SSE2/ASIMD x4 or AVX2 x8), AVX2 emits packed FMA when fast math is enabled, and AVX-width kernels emitvzeroupperon return. AVX-512 is detected but remains disabled until a measured microarchitecture cost model is available.YdspBenchmarkTestsnow compare automatic code against the scalar baseline on the modal-bank shape, reporting lane width, generated code size and timing while requiring identical strict output. -
YDSP: the vectoriser now handles the scalar remainder of a trip count, and the per-sample stream loop. A constant-bound loop whose span is not a whole multiple of the lane count peels its leading remainder as straight-line scalar copies and starts the vector loop at a whole number of vectors (a 6-mode bank is two scalar iterations plus one four-lane trip, not six scalar ones); a non-zero constant start is handled the same way over the span, which also fixes an overrun the old divisible-bound-only rule had for such loops. A
blockSize-bound loop whose body reads or writes streams at the loop variable (in[i]/out[i], the sample-mode gain/mix shape or a block-mode stream loop) is widened too, with a rolled scalar tail loop after the vector loop, packed stream loads/stores on the native backends andv128.load/v128.storeon wasm SIMD, and the reduction fold placed once after the tail. Loops that were already scalar stay scalar: a constant span shorter than one vector, a runtime start, ablockSize-boundstate-array bank, or a stream store at a fixed index. -
YDSP: missed-vectorization diagnostics.
YdspVectorizer::runnow records one outcome per original loop -widenedat the lane count, or the exact reason the loop stayed scalar (shortTripCount,unsupportedWidenedOp,indirectAccess,loopCarriedValue,nonConstantStart, ...) - exposed asYdspKernelReport::loopVectorization(withYdspVectorizationReport::rejectionReasons()for the deduplicated text andYdspVectorizationResult::describe()for one line), and emitted as info diagnostics whenYdspCompileOptions::emitOptimizationReportis set, the LLVM-Rpassequivalent for "why is this loop scalar?". -
YDSP: the WebAssembly backend now lowers the vectorised IR to f32x4 SIMD when the module is compiled with
-msimd128(the emscripten default, which defines__wasm_simd128__). Widened state-bank loops emitv128.load/v128.store, packedf32x4arithmetic, splats and a shuffle-based horizontal reduction; element-wise work stays bit-exact against the scalar form and the reassociated accumulation keeps the existing tolerance. The automatic tier applies the full native transform set on wasm: vectorisation at four f32x4 lanes, unrolling of the widened loops, and halving of the widened reduction chains (all pure IR passes), while a build without-msimd128keeps every loop transform off and the scalar-only rejection, now naming the flag. The wasm vector width is fixed at four lanes (128-bit SIMD), so kernels vectorized for AVX2/AVX-512 widths are rejected with a diagnostic rather than miscompiled. -
MIDI: the WASM (
Emscripten) backend ofyup_audio_devicesis now backed by the Web MIDI API (yup_Midi_wasm.cpp).MidiInputandMidiOutputenumerate, open, start/stop and send to browser MIDI ports; SysEx is requested (sysex: true); hot-plugstatechangeevents updateMidiDeviceListConnectionlisteners; incoming streams are converted through the existing bytestream handlers (soump::ReceiverwithMIDI_2_0works), outgoingump::View/ump::Packetsare converted to MIDI 1.0 bytes, and sends from other threads are proxied to the main thread.createNewDevice()stays unsupported since the Web MIDI API cannot create virtual ports. Note the browser permission is asynchronous: callgetAvailableDevices()early and lists populate once the user grants access. -
Tools: a new
modules/yup_dsp_jit/tools/vscode-ydspextension brings YDSP syntax highlighting, snippets and editing configuration to VSCode, installed withjust vscode. -
Examples:
examples/graphicscan now be configured to build a single demo instead of the full browser, via-DYUP_EXAMPLE_GRAPHICS_DEMO=<id>(e.g.SpinningCube) in place of the defaultALL. A single-demo build compiles out every other demo's code, skips the picker list UI, and only embeds/preloads the rive, lottie, shader anddata/synthsresources (and theglslang/spirv_cross/spirv_toolsshader transpiler) that the selected demo actually needs - useful for small, single-page embeds such as documentation website demos. -
GUI:
yup_audio_guigainsPitchWheelComponentandModWheelComponent, vertical-drag controls for building a pitch-bend + mod wheel strip next to aMidiKeyboardComponent. The pitch wheel reports a bipolar-1.0..1.0value and springs back to its default on mouse release by default (settable to hold instead); the mod wheel reports a unipolar0.0..1.0value and never springs back. Both are themed as a cylindrical wheel body with a single sliding grip line - not a slider track and thumb.examples/graphics's YDSP Synth Lab now drives pitch bend and CC1 from the two wheels, placed to the left of its keyboard, instead of the two sliders that previously stood in for them in the expression row. -
YDSP:
noteOnhandlers can now reade.bendSemitones, the pitch-bend in effect at the moment the note is triggered. Previously a freshly triggered voice only saw the bend as a laterpitchBendevent, so patches either reset their bend factor to 1.0 on note-on (a key pressed while the wheel was held up started un-bent) or leaked a stale value from a recycled voice. The runtime already computed the value (payload.bendfrom the MPE note, which in legacy mode is the last wheel position on the channel); it is now exposed on thenoteOnshape likepitch/velocity, and mono note-ons carry it too (the mono held-note record now stores the bend). The bundled synth patches replace theirbendFactor = 1.0note-on reset withbendFactor = pow (2.0, e.bendSemitones / 12.0). -
YDSP: processors can now generate events.
output event <name>;declares an emitting channel andemit <shape> (field: expr, ...) -> <name>;sends one, legal in the per-sampleprocessbody and in event handlers. A channel's declared name is an identifier, not a shape - it need not equal any of the seven shape names, and one channel may carry several different shapes over its lifetime. In aconnection { }block, a node'soutput eventwires to another node'sinput event(including the polymorphicmidiinput) or to the graph's ownoutput eventboundary, which the host reads back as MIDI; every declaredoutput eventmust be connected at least once, and an emitted event straddling a block boundary or carrying[[ latency ]]compensation arrives already aligned with its source node's audio. This is what lets a MIDI-only processor - one with no stream endpoints at all, such as the newmidi.Arpandmidi.Transposeinexamples/graphics/data/synths/midi/- drive an existing, unmodified voice bank by composition alone;ArpPolySine.ydspandArpTranspose.ydspare worked examples, the latter a fully MIDI-only graph with no audio stream anywhere in the patch. A graph is now classified purely by which endpoint kinds it declares (audio-only, MIDI-only, or hybrid), not by a keyword. A MIDI-only node's own note bookkeeping (e.g.midi.Arp's held-note table) is no longer subject to the runtime's ordinary per-voice allocation and stealing: with no stream endpoints, and therefore no per-note "sound" to steal, every event now reaches the node's sole instance directly, so a node likemidi.Arpcorrectly tracks a whole chord rather than just its most recent note.midi.Arpalso gained amodeparameter (Up,Down,Up-Down) selecting which direction it steps through the held notes, and restarts its clock on the first note of a phrase so a chord sounds immediately rather than after up to one full1/rate. -
YDSP: fixed a bug in the compiler's import-cloning step (
cloneStmt()) that dropped anemitstatement's shape, target and field list whenever the emitting processor was only ever reached through animport- every existing test that usedemitdeclared its processor inline, so the gap went unnoticed untilmidi.Arp/midi.Transposeshipped as importable library processors. -
YDSP: a graph's
input event <name>;is no longer a broadcast subscription. Previously any node whose own processor declaredinput event <name>;with the same name received every event on that port automatically, with no wire in theconnection { }block to show it - and which port a node actually heard depended on its declaration order among the graph's inputs, invisible at the call site. A graph input event is now wired exactly like a stream:<name> -> node.event;in aconnection { }block, fanning out to as many destinations as are wired and reaching none that aren't; an unconnected graph input event or an unconnected node input event is now a compile error, the same rule already applied tooutput event. All shipped synth patches (andArpPolySine.ydsp/ArpTranspose.ydsp) gained the explicitmidi -> voices.midi;(ormidiIn -> arp.midiIn;) line this requires;PolySine.ydspmoved from the algebra body form to aconnection { }block, since the algebra form has no event syntax. -
YDSP: native transcendentals now lower through the bundled
sleef_library(SLEEF, Boost-licensed): widened transcendental loops vectorize to 4-lane calls (u35under the now-defaultfastMath,u10when strict), scalar float32 values stay on libm,fastMathis enabled by default on native targets (wasm stays strict regardless), and an 8-lane AVX2 value splits into two 4-lane calls. Optimizer borrows from the SNEX reference JIT: constant division becomes reciprocal multiplication under fastMath, pow2 modulo of a provably non-negative value becomes a mask, and adjacent same-bound memory-disjoint loops fuse before vectorization; constant math calls fold for the full intrinsic family. Register allocation weights hot values (inductions, widened lanes, stream bases) so they never spill. Benchmarks report a per-policy matrix (baseline strict / + fastMath / host strict / host + fastMath default) with a transcendental-call count per kernel. -
YDSP native codegen: under the native
fastMathdefault, scalar float32expon AArch64 is lowered to a straight-line degree-8 Estrin polynomial (~1 ulp across [-1, 1]) instead of a per-sample libm call plus its register-allocator spill round-trip; arguments beyond |x| = 1 take a rare, predictable fallback branch to libm, and x86-64 keeps the libm call. The exp-envelope benchmark drops from ~4.6 to ~1.8 ns/sample (~1.05x of its C++ reference, down from 2.6x) and the modal bank from ~7.3 to ~5.0 ns/sample (0.74x of C++). The optimizer's block-local CSE now also covers repeated reads of the same input stream slot - input buffers are immutable for the lifetime of a kernel, so the reads are pure - collapsing e.g. the four per-samplein[i]loads of the delay-taps shape to one. -
YDSP optimizer: scalar state write-backs fold into the value they move. A sample-mode state update lowers to
v = op (...)followed bymovF y = v(and the builder keeps binding the variable tov); the newfoldStateWriteBackspass rewrites the producer to write the loop-carried registerydirectly, redirects every in-block use ofvto it and drops the move, so the per-sample chain carries no extra register hop. The ladder filter drops from ~10.7 to ~9.7 ns/sample (host + fastMath, now ~1.05x of its C++ reference), the wave folder from ~1.8 to ~1.1 ns/sample (~1.03x of C++), and the exp envelope from ~1.8 to ~1.5 (0.86x).
-
YDSP: the lexer scans with a UTF-8 character pointer instead of a character index into the source
String.String::operator[],length()andsubstring()each walk the buffer from the head, and the old cursor called them once or twice per source character, which madeYdspLexer::tokenize()quadratic in source length; it is now linear. Tokens, line numbers and column numbers are unchanged - columns still count characters, not bytes. -
Tools: the python stdlib archive generator (used by the tests target) now skips the copy and zip steps when the source bundle and tool configuration are unchanged since the last run, so reconfigures no longer pay for a full stdlib rebuild.
-
Build:
yup_add_embedded_binary_resourcesnow regenerates a resource's byte array only when the input file's content actually changed (tracked via an MD5 sidecar next to the generated.inc), so reconfigures no longer re-read and re-serialize large embedded files such as the python stdlib zip. -
Build: Xcode builds no longer auto-regenerate the project during a build (
CMAKE_SUPPRESS_REGENERATIONis set for the Xcode generator), so building no longer cancels with "project is being modified while building" when CMake input files change; re-run cmake (e.g.just mac) after editingCMakeLists.txtor.cmakefiles. -
YDSP native codegen: kernel prologues now load context pointers only when the generated IR uses the corresponding resource, reducing register pressure and avoidable spills in small sample kernels.
-
Emscripten: the standalone shell now shows a non-blocking hint over the canvas when audio needs a user gesture, reports audio/MIDI availability in the top rail, and keeps activation in the shared AudioWorklet backend for all examples using the shell.
-
GUI:
PitchWheelComponentandModWheelComponentnow take aMidiKeyboardStatein their constructor, likeMidiKeyboardComponent. Both register as state listeners and follow the pitch-wheel / modulation-wheel (CC 1) position of their midi channel (seesetMidiChannel()), applied asynchronously on the message thread - updates arriving between message-thread passes are coalesced so only the latest position is applied, and no update is applied while the user is dragging the wheel. The YDSP Synth Lab demo drops its manual atomic bridge (incoming pitch bend / CC1 stored on the MIDI input thread and applied inrefreshDisplay()) and lets the wheels followkeyboardStatedirectly. -
YDSP optimizer: the compiler pass implementations are split out of the monolithic
optimiser/passes/yup_YdspPasses.cppinto one file per pass (yup_YdspPassesConstantFolding.cpp,yup_YdspPassesAlgebraicSimplification.cpp,yup_YdspPassesCopyPropagation.cpp,yup_YdspPassesIfConversion.cpp,yup_YdspPassesFullyUnrollBoundedLoops.cpp,yup_YdspPassesSplitWidenedReductionChains.cpp,yup_YdspPassesStoreToLoadForwarding.cpp,yup_YdspPassesDeadCodeElimination.cpp,yup_YdspPassesLoopInvariantCodeMotion.cpp,yup_YdspPassesContractMultiplyAdd.cpp,yup_YdspPassesLowerFusedMultiplyAdd.cpp), with the helpers shared across passes moved toyup_YdspPassesShared.cpp. Pure cut/paste - no behavior or IR change. -
YDSP: the public API headers are split out of the monolithic
compiler/yup_YdspCompiler.hinto per-concern files: the runtime graph API (YdspAudioGraph,YdspParameterInfo,YdspExecutionReport, the stream buffers andYdspProcessResult) now lives inruntime/(yup_YdspAudioGraph.h,yup_YdspTypes.h,yup_YdspExecutionReport.h), and the compiler API (YdspCompiler,YdspDiagnostics,YdspCompileOptions/YdspOptimizationReport,YdspRecursionGuard) incompiler/(yup_YdspCompiler.h,yup_YdspDiagnostics.h/.cpp,yup_YdspCompileOptions.h,yup_YdspRecursionGuard.h). Pure cut/paste - no behavior or API change. -
Examples: the YDSP Synth Lab demo (
examples/graphics/source/examples/YdspSynths.h) is restyled to matchcmake/platforms/emscripten/shell.html's dark theme - the same void/surface/edge/ink/muted/glow palette, and a surface-coloured rail with a hairline edge behind the toolbar/title bars and behind the keyboard row. The toolbar is now one aligned strip: the YUP mark and a bold "YUP!" wordmark, a shortened patch selector, then Performance/Editor tabs, master volume, the oscilloscope and All Notes Off as equal-width slots, so the row reads as a single set of controls rather than mismatched widths; the separate "YDSP Synth Lab" title label was dropped as a redundant second wordmark. The Dump Asm/Dump Wasm button moved out of the toolbar into the editor tab, next to Compile, since it only makes sense there. The row below the toolbar now holds the MIDI input selector alongside the expression sliders, reclaiming the space the old dedicated title/patch/volume row used to waste; and the parameter knob grid expanded into that freed height plus the space the oscilloscope vacated. Parameter cards, meters and the oscilloscope now share the shell's blue accent instead of each having their own.data/logo.pngis preloaded unconditionally in the Emscripten build, sincemain.cpp's own title chrome loads it regardless of which demo is selected and it previously was not. -
Emscripten: the standalone shell page (
cmake/platforms/emscripten/shell.html) has been redesigned around a dark, low-chrome theme keyed to the YUP mark, with the loading indicator doubling as the download progress arc. The canvas display is now a three-way choice - embedded, full window (canvas fills the tab, no scrollbars) and fullscreen - replacing the old resize-canvas and lock-pointer checkboxes, which only fedModule.requestFullscreenand had no observable effect. Fullscreen now reuses the full window layout rather than emscripten's own fullscreen sizing, so the two behave identically. The top rail reports whether the page is cross-origin isolated and how many threads the browser offers, which is what a failing pthreads or audio worklet build needs first. Note that full window scales the existing framebuffer: the app is only redrawn at the new size onceSDL_EVENT_WINDOW_RESIZEDis forwarded tohandleResizedinyup_Windowing_sdl.cpp. -
YDSP: new
fma(a,b,c)intrinsic -a * b + cwith a single rounding. The compiler has never fused a multiply and an add on its own, deliberately: contraction is a precision liberty, and a patch has to produce identical samples on every backend it can be compiled for. That refusal turned out to be most of what separates the JIT from compiled C++ on per-sample recurrences, where the multiply and the add are consecutive links of the loop-carried chain. A new benchmark variant sizes it by giving the C++ reference a#pragma clang fp contract(off): the ladder filter reads 1.56x against a reference that fuses and 1.05x against one that cannot, so on that shape contraction was essentially the entire gap; the wave folder reads 2.15x and 1.24x. The wave shaper is the control - its recurrence is a multiply and aselect, with no add to fuse into - and its two references land together, confirming its ~1.39x is not an FMA story.fmacloses the gap without giving up the guarantee: the operation has one defined value, and a target with no fused instruction (wasm, or x86-64 without FMA3) reaches that same value by computing in float64 and rounding once - exact for normal results, since a float32 product is exact in float64 and 2p+2 = 50 bits fit in its 53. So native and wasm still agree; what differs is a patch written withfmaagainst the same patch written with*and+. It is float32 only, because the fallback needs a format one step wider than the operands and none exists above float64 -fmaon float64 operands is a compile error rather than a silent per-target difference. The expansion is an IR pass, so neither the wasm backend nor a pre-FMA3 x86 one needs to know the opcode exists; AArch64 lowers it tofmadd, x86-64 with FMA3 tovfmadd213ss. AcontractMultiplyAddpass applies the same rewrite automatically to everya * b + cwhose multiply feeds nothing else, so patches get this without being rewritten - and measurably do: with it on, the ladder went 15.2 to 10.3 ns/sample (1.56x to 1.135x) and the wave folder 2.71 to 1.87 (2.15x to 1.645x), both now faster than the reference that cannot fuse (0.77x and 0.94x). The same patches written withfma()by hand measure 0.98x and 0.99x against the automatic form - inside noise, which is the check that the pass finds what a person would and picks the same operand when both are eligible. The wave shaper is the control: its recurrence is a multiply and aselectwith no add to fuse into, and it did not move (1.39x to 1.41x). The pass runs after the vectoriser and skips anything widened, since there is no portable packed fused form; when both operands of an add are fusable multiplies it picks the one on the recurrence, because fusing the other leaves the loop-carried chain a link longer than it started. Two costs come with it being automatic: on a target with no fused instruction each contracted site is six operations instead of two, which WebAssembly pays throughout and which is not yet measured; and an algorithm that depends on the intermediate product being rounded (a*b - c*ddeterminants, Kahan summation, Dekker'stwoProduct) is changed by it, with no per-expression way to opt out yet. -
YDSP: a graph is now an arbitrary DAG. Connectivity was "exactly once" on every graph input, graph output, node input and node output stream, which made a YDSP graph a forest of chains: a signal could not be split and rejoined, so dry/wet, parallel multiband, mid/side, a metering tap and summing two sources into one input were all inexpressible. The rule is now at least once on both sides, and beyond that: a source may fan out to any number of destinations, and a destination fed by more than one source sums them, with no mixer node needed. Fan-out costs nothing - generated code never writes through
ctx.inputs, so N consumers share one buffer pointer and no new memory is allocated. A newYdspBenchmarkTestsshape measures the cost, and the answer is that summing is effectively free. A graph-level dry/wet built from fan-out plus fan-in runs at 6.32 ns/sample against 5.96 for the same patch with dry/wet hand-rolled inside one processor - but most of that 0.43 gap is the second kernel call, not the mix, because the fanned form cannot fuse. A third variant in the same test separates them by fanning out to two separate graph outputs instead of summing into one: same two kernels, same fan-out, but both outputs take the direct-write path so no mix buffer exists. Against that, the mix path costs 0.109 ns/sample (6.32 vs 6.21, a 1.018x ratio) - amemcpyplus one add per sample, which is what it should be. Graph-level dry/wet is therefore the idiomatic form now rather than a luxury, and the remaining ~0.32 ns/sample is simply what a second kernel call costs. Implicit summing requires afloat32orfloat64stream, named in the diagnostic when it is not; fan-out has no type restriction, being the same buffer read twice. The four zero-use rejections survive and two of them are load-bearing rather than stylistic: a graph output with no source would have nothing written to it (and the runtime never zeroes an output buffer, so the host would hear its own uninitialised memory, not silence), and a node input with no connection leavesruntimeInputs[s]null for the kernel to dereference. Mixing happens on the input side and the destination owns the mix buffer, which leaves a node's exclusive ownership of itsruntimeOutputs- and therefore the polyphonic pre-voice-loop zeroing,clearVoiceSpanand the voice accumulate - untouched; connection 0memcpys into the buffer and 1..N-1 accumulate, so "written exactly once" is structural rather than an ordering rule to get right. Summation order is the order the edges appear after analysis: deterministic per patch, and nothing more (subgraph inlining and kernel fusion both rebuild the edge list), so it is preserved with a counting sort rather than astd::sort, and documented as something not to depend on. Existing patches are bit-identical: a graph output keeps writing straight into the host buffer when it has exactly one source, that source is a node, the edge carries no delay, and that node output has exactly one destination - which every previously-legal shape satisfies.MasterBus.ydspis the first patch in the tree to use the feature, gaining a real parallel dry path around its reverb soreverbMixis a graph-level balance instead of a value forwarded intoReverb's own internalmix; three new demo patches showcase it:ParallelRack.ydsp(one voice fanned out to three character paths, summed back),HaasWidener.ydsp(one mono chain fanned out to both graph outputs with an inline delay on the right), andParallelDrive.ydsp(a clean path summed against an oversampled hard clipper feeding a[[ latency: 32 ]]lookahead limiter, so the 48-sample skew is compensated rather than comb-filtering the blend). The four bundled effects are deliberately left alone as regression baselines. -
YDSP: automatic plugin delay compensation, and a latency figure to report to the host. Before paths could reconverge, latency misalignment could not be heard, so YDSP had no latency concept at all - zero occurrences of the word in the module. The moment fan-in exists, an oversampled branch summed against a dry one is comb-filtered rather than merely late. The new
computeLatencyAndCompensatepass equalises it, andYdspAudioGraph::getLatencySamples()reports what is left (feed it toAudioProcessorBase::setLatencySamples(), which already drives every VST3/CLAP/AU/AUv3/AAX/LV2 wrapper). What makes it correct rather than merely present is which latency it compensates: an artifact, where the sample count is a leaked consequence of an implementation choice the author did not make (an oversampler's group delay; a newprocessor P [[ latency: N ]]declaration, which only the processor's author can know), as opposed to intentional, where the count is the semantics because the author typed the number (-> [400] ->,x @ 400).edge.delaySamplesappears nowhere in the model. The decisive case is a dry/wet delay effect: the wet path's@is intentional, so both branches have artifact latency 0, nothing is inserted, the dry stays dry, and the patch reports 0 - where a naive "total group delay" model would delay the dry branch, destroying the effect, and tell the host a 500 ms echo was 500 ms of plugin latency. Compensation always lands on the reconvergence edge, never hoisted upstream, so a branch with one incoming edge always gets 0 and a plain chain is untouched. Separate graph outputs are equalised against each other as well, which is forced rather than chosen: every format YUP targets reports one scalar, so a per-output latency vector is unrepresentable and a patch whose L is 16 samples later than its R would be permanently skewed in every host with no diagnostic - while a wanted skew, written-> [16] ->, is never touched. The oversampler's contribution is derived from the sameydspOversamplerSincRadiusconstant the runtime instantiates itsyup::Oversamplerwith (16 input-rate samples, linear phase, exactly integral for every factor - which is why an integer delay compensates it perfectly) rather than restated in a comment.[[ latency: N ]]is declared in the processor's own sample domain, so a* 4instance divides by 4 and a non-dividing factor is a hard error naming both numbers; rounding would ship a sub-sample residual inside the one feature whose job is phase alignment.YdspAnalyzedEdge::compensationSamplesis kept apart fromdelaySamplesso the latter keeps meaning "what the author wrote", the fusion predicate keeps its meaning and the pass stays idempotent - the compiler sums them at one line.YdspAnalyzedNode::latencySamplesis the sole source of truth afterwards, becausefuseNodeChainssynthesises a processor declaration that would report 0; a fused node's latency is set to the sum of its members'. The pass runs after fusion, so it cannot cost a fusion opportunity. The reported value is a compile-time constant: a YDSP graph is a fixed DAG, so nothing at runtime can reroute it or change an oversampling factor, and a plugin offering an oversampling selector recompiles on the control thread instead - which is also all the formats support, since changing latency needs a restart request rather than a realtime notification. -
YDSP: the
<:(split) and:>(merge) algebra operators are real. Both were tokenised, given their ownYdspOperator, parsed at one precedence level with:and then quietly handed tocomposeSequential- so they behaved as a plain:and the documented arity rules were never applied (and were stated backwards:<:widens,:>narrows).AlgebraValue's port lists became bundles - one set of terminals per channel rather than one terminal - because "channel 0 goes to two places" and "channel 0 is the sum of two places" cannot be said otherwise, and every wire-emitting site became a cross product over them. All three operators now share onecomposeFanned:a <: brequiresa.outArityto divideb.inArityand assignsb's input j toa's output j % a.outArity;a :> brequiresb.inArityto dividea.outArity, sendsa's output i tob's input i % b.inArity, and sums the collisions;:is the case where the two are equal. The result arity is(a.inArity, b.outArity)throughout. Neither operator needs a relay node, and neither does_: it carries no ports, so_ <: (a , b)and(a , b) :> _emit zero wires and are handled by regrouping the other operand's bundles, with the fan materialising when the value is later sequenced against a real leaf. An unconstrained_defaults to arity 1 on the side the operator governs (left of<:, right of:>- the maximum fan) and to the known side's arity otherwise.process = dry <: (Distort , Chorus) :> wet;is now a parallel dry/wet, and compiles to a graph bit-identical to writing its four edges out by hand. This only became implementable once fan-out and summing fan-in existed: a split is a fanned-out edge and a merge is a summed one. -
YDSP: a subgraph boundary port can now carry a fan on either side.
inlineSubgraphsheld oneintper boundary port, written unconditionally, so a second edge ontosub.in(or out ofsub.out) silently replaced the first and one source or destination vanished - with fan-in now analyzable, that would have been silent wrong audio rather than a compile error. Each side is a list, and the splice is the cross product of the sources reaching a producer's boundary and the destinations leaving a consumer's boundary; the nested loop is required for a pass-through edge inside a subgraph, which touches both boundaries at once and which the two previous independentifs only handled for the 1x1 case. Each resulting edge accumulates the internal edge's own delay plus the parent edge's delay at each end, added per resulting edge rather than hoisted out of the loop.paramAlias/meterAliasstay single ints: a node parameter still takes one writer. -
YDSP: undersampling (
node x = P / 4) is a real decimator. It previously only shortenedctx.numSamples, so the kernel filledblockSize / Nsamples and the rest of the block kept whatever the previous one left there, while an odd block size dropped the remainder outright - it was briefly rejected outright during this work rather than shipped broken. It now band-limits and decimates the node's inputs, runs the kernel at 1/N, and interpolates its outputs back, reusing the sameyup::Oversamplerwindowed-sinc machinery*Ndoes (factors 2, 4 and 8;float32streams only, like*N). TwoOversamplerinstances are needed rather than one, driven from opposite ends:upsample()resizes the shared oversampled buffer, so a single instance's direct writes for decimation would fight with it. That in turn neededdownsample()to stop requiring a precedingupsample()- it now derives the oversampled length from its ownnumSamples * OversampleFactor(which the assert it replaced already proved equal) and its capacity checks are against whatprepare()allocated, so the class is usable in either direction with no new API.getOversampledChannelData()is bounded by the allocation for the same reason: a decimate-first caller has to fill that buffer, and would otherwise have nowhere to write. "Is a block pending" remainsgetOversampledNumSamples()'s job, which is the accessor designed for it. Four newOversamplerTestcases cover the direction (accessor availability afterprepare(), standalone anti-aliased decimation, a decimate-then-interpolate round trip in the two-instance shape, and continuity across block boundaries). The block-size problem*Ndoes not have -/Ncan only consume whole groups of N samples, so any block size that is not a multiple of N leaves a remainder - is solved with two carry FIFOs whose counts are invariant atN - 1between them, which is exactly the surplus needed to fill every block; priming the output side withN - 1zeros is what establishes it, and costsN - 1samples of latency. Total/Nlatency is16 * N + (N - 1)graph-rate samples (the resampler's2 * SincRadiusis in the node's domain, hence the multiplication), all of it artifact latency that the new delay compensation removes automatically - which is exactly what the artifact/intentional rule predicted for a decimator before one existed. Relatedly,sampleRateinside a rate-changed kernel now reports the rate that kernel is actually running at rather than the graph's; it was previously always the graph's, so an oversampled processor deriving any coefficient from it (which is the only reason to ask) ran those coefficients N times too slow, and an undersampled one would have run them N times too fast. The newControlRateWah.ydspdemo is built on both: an envelope follower - control-rate work by nature - runs at/ 8while the audio path stays full-rate, and the 135-sample skew where the two reconverge is compensated away, which is what stops the filter opening ~2.8 ms after the transient that opened it. -
YDSP: the feedback-cycle diagnostic no longer implies an inline delay would fix it.
rebuildTopoOrderrejects every cycle but said "feedback cycle without a delay", which read as a promise that-> [1] ->would break one. It now says plainly that cycles are unsupported in this version and that a delay does not break them. -
YDSP (breaking): four syntax changes. Function parameters are now
name: type-func noteToFreq (pitch: float) : float { ... }- and the old(float pitch)form no longer parses. Imports use dotted module paths:import fx.Delaymaps to the filefx/Delay.ydspand forces access asDelay.*(import X.Y.Z as WforcesW.*); importing two different files that would share a namespace is a compile error suggesting anasalias, and a circular import is now a hard error (it was a warning). A file that declares only top-levelfuncs is a library, imported like anything else and called asns.funcName (...). Events are now named channels:input event <name>;at processor and graph scope, any name and any number (like streams), each carrying all seven shapes; a handler selects the shape withevent midi (e: noteOn) { ... }(the old per-shapeinput event noteOn;/event noteOn (e)pair is gone). Each graph event input is a separate event stream - a node subscribes to the graph input matching its processor'sinput eventname, andprocess()gained a span-based overload (yup::Span<const yup::MidiBuffer*>) plusgetEventInputCount()/getEventInputName()to feed them individually; the single-buffer overload feeds the first input. -
YDSP:
YdspCompiler::compile()takes an optionalThreadPool*(third parameter). When one is passed, imported files are read, lexed and parsed in parallel on that pool; the merge stays single-threaded, the results are identical to the sequential path, and the pool remains caller-owned - the compiler never removes jobs it did not add. -
YDSP: a graph parameter can now drive more than one node parameter. It was capped at one, which made a wrapper graph unable to do the one thing it exists for -
MasterBuscould not forward a singlemixto both its compressor and its reverb, and the diagnostic read as an arbitrary limit. The cap is gone fromvalidateConnectivity(a node parameter still takes at most one writer, which is what keeps the aliasing unambiguous), andPimpl::paramSlotToNodebecame a list per slot so sample-accurate automation is queued on every node the slot drives instead of only the last one wired - previously a second edge silently overwrote the first's routing, so automating that parameter moved one stage and left the other at its old value. The per-block copy path needed nothing:node.paramCopieswas already per-node. -
YDSP: graphs can now be composed. A program may declare several
graphblocks - the entry point is the one annotatedgraph Name [[ main ]] { ... }, or, when the program declares just one graph of its own, that one - and anodemay instantiate a graph instead of a processor, so a reusable chunk of signal flow (an effects chain, a voice architecture) gets a name and a parameter surface of its own.importnow merges the imported file's graphs alongside its processors under the same namespace prefix, sonode master = fx.MasterBus;works across files; an imported graph is always a subgraph, whatever[[ main ]]it carries for its own file. A subgraph is compiled away: the semantic analyzer splices its nodes and edges into the graph that uses it, so the optimiser, all three backends and the whole runtime keep seeing one flat graph and nothing past analysis changed. Inner nodes take the instance path as a name prefix (master.compressor.threshold), a subgraph'sinput valueforwards to whatever node parameter it drives (the node's override list sets it; a parameter edge onto the subgraph node aliases straight through), anoutput valueforwards outwards the same way, and inline delays on both sides of the boundary add up into one edge. Graphs are inlined innermost-first in dependency order, so nesting is unbounded and a graph that reaches itself is a compile error naming the loop. The three things that are runtime mechanisms rather than wiring are rejected on a subgraph node with an explicit reason: a voice bank (Sub[16]) needs one processor with one float32 output stream for the per-voice summing path, over/undersampling (Sub * 4) is a per-processor rate change, and graph-scopeinput event midibelongs to the entry point. A subgraph parameter that drives no node inside it warns rather than silently doing nothing. The demo's "Analog Saw" patch uses the newexamples/graphics/data/synths/fx/MasterBus.ydsp, a graph composing the existingCompressorandReverbprocessors. -
YDSP: a voice bank can now skip idle voices. A processor nominates one
state intscalar as its activity flag with the new state annotationstate int active [[ role: voiceActivity ]];, and the runtime skips a voice whose flag reads 0 and whose key is not held - no kernel call, no output accumulation - so aVoice[16]bank playing a three-note chord pays for three voices instead of sixteen, and an idle one pays for none. The patch declares completion itself (it knows its own delay lines, high-Q filters and release curves) by testing the amplitude it actually emits; the runtime supplies theheldcheck, so a patch that clears its flag too early wastes CPU on a held voice rather than going silently mute. The flag is read once per voice per block and re-checked after an event handler runs, so a mid-block note-on wakes the voice for the rest of the block, and a voice is never put back to sleep mid-block. Entirely opt-in: a processor that declares no flag keeps byte-identical behaviour. An unknownrole:value, a non-intor array-typed flag, a second flag, an initialiser, or the annotation on a processor with no event handlers are all compile errors. The shippedPolySine,AnalogSaw,FMBell,WobbleLead,PulseBassandElectricPianopatches opt in. NewYdspAudioGraph::getActiveVoiceCount ("nodeName")reports how many voices will run on the next block, sharing the scheduler's predicate exactly. -
YDSP: new kernel-fusion pass. A chain of nodes that only feed each other -
x : Osc : Filter : Gain : y, or the same wiring as aconnectionblock - is compiled into one kernel instead of three, so the twoblockSizescratch buffers the intermediates round-tripped through become registers. Measured on a newYdspBenchmarkTestsshape that compares a three-stage chain against a single processor computing the same thing: the chain went from 5.68 to 2.68 ns/sample, a 2.1x speedup, landing on the hand-fused processor's 2.70 ns. The shape had measured a 2.05x gap before the pass existed - the win had been asserted from first principles and never measured, which is why the item stayed open - so the ratio now sitting at ~1.0 is what the shape guards. The pass works off the analyzed graph, so the:algebra and an explicitconnectionblock fuse alike, and it runs after subgraph inlining, so a chain assembled out of a subgraph's nodes fuses exactly like a hand-written one. A link is fused only when the producer feeds nothing else and the consumer is fed by nothing else - an intermediate tapped to a second destination is observable and stays a real buffer - the connection carries no inline delay ([N], which only the runtime's delay buffer can provide), and both ends are per-sample single-in/single-out kernels that are not voice banks, rate-changed or event-driven. The fused processor is synthesized as an ordinaryprocessorand analyzed through the normal path, so the IR builder, all three codegens and the runtime scheduler see nothing new; every name a member declares is rewritten with a per-member prefix, since ablockstatement does not open a scope in this language and two members' locals would otherwise collide. Host-visible parameter and meter names are preserved -first.gainandfirst.levelstill resolve after the node they belonged to has been fused away - via a new per-node public-name override, a graph parameter aliased onto a member's parameter still drives it, and a member's meter wired to a graphoutput valueis rerouted onto the fused node so it keeps reporting. Once absorbed, a member's own kernel is dropped rather than JIT-compiled as unreachable machine code thatgetExecutionReport()would list as something the patch runs; only what fusion orphaned is dropped, so a processor still instantiated on another branch survives.YdspDiagnosticsgainedmark()/rollbackTo()so that a synthesized processor which fails to analyze is silently declined rather than failing the user's compile against source they never wrote. -
YDSP optimizer: new bounded-loop vectoriser. On the native backends a constant-bound loop over parallel
state floatarrays - an oscillator, partial or modal bank - is widened to four float32 lanes, so a 16-mode loop runs four packed iterations instead of sixteen scalar ones. It is an IR → IR pass rather than codegen widening, so both native targets share it and it is testable without a JIT; a widened value simply carrieslanes > 1and the existing arithmetic opcodes are reused (addFat four lanes isaddps/fadd v.4s), with onlyvsplatandvreduceAddFadded for the two places where the operand and result lane counts differ. Only trip counts that are a whole multiple of four are widened, so there is no scalar epilogue and no block is created, emptied or reordered - the CFG-linear layout both backends recover regions from is preserved by construction. The pass runs last, after loop-invariant code motion has hoisted the loop's shared work (a bank'sexp-derived drive term) into the preheader, where it stays scalar and is broadcast once. An accumulation (sum = sum + z[i]) becomes a vector accumulator folded once on the way out, which is what breaks the serial dependency between iterations - and it reassociates, so a widened reduction is not bit-exact while element-wise widening is. A widened body also sinks its loop-invariant scalar prelude into the preheader:inread inside a loop lands there as aloadInputand stays, because loop-invariant code motion never hoists a load, but in a widened body it provably cannot be invalidated -storeOutputdisqualifies the loop outright, so nothing in it can write stream memory - and the index is already known to be loop-invariant. The load, everything pure that depends only on it, and the broadcast they feed therefore run once per loop entry instead of once per element. State-array loads are excluded:z[j]at a loop-invariantjis the same element as the widenedz[i]store on whichever iterationi == j. Measured against the hand-written C++ equivalents inYdspBenchmarkTestson AArch64: the 32-partial harmonic bank went from 3.20x slower to 0.82x - faster than the compiled reference - and the 16-mode modal bank from 7.51x to 2.35x, of which the last 0.65x came from that hoist. Shapes with no qualifying loop - a delay line, a ladder filter, a wave shaper - are unchanged. A loop is left scalar unless every array element is reached through the loop variable itself, the body is a single block, the loop variable is used for nothing but indexing, and nothing in it touches a stream, a scalarstateslot, a'/@/smoothslot, a transcendental, a comparison or aselecton a widened value.YdspKernelReportgainedvectorizedandvectorWidth. The wasm backend has no0xFD-prefix opcode family, so that path stays scalar and rejects a widened kernel outright rather than emitting scalar code for packed values. -
YDSP optimizer: new
ifConversionpass. A short, else-lessifwhose body is a single side-effect-free block becomes straight-line code plus oneselectper assignment, turning a data-dependent branch in a sample loop into a conditional move - a wavefolder taking two such branches per sample was paying an unpredictable branch for each while compiled code emittedfcsel. The body is only speculated when every instruction can be executed unconditionally: no memory access (a guarded index may be out of range when the guard is false), no call, and at most eight instructions, so both paths together stay cheaper than a mispredict. The assignments keep their original order, so an instruction that reads a register the body already assigned still sees the conditionally-updated value. No block is added or removed - the condition block absorbs the body and both fall through - so every block index, loop bound and wasm structured region is unchanged. -
YDSP optimizer:
loopInvariantCodeMotionnow hoists into each loop's own preheader instead of only the function entry block, using real dominator sets rather than an "is it in block 0" test. Work that is invariant across an inner loop but varies per sample - a filter coefficient derived from an envelope, say - previously could not be hoisted anywhere: the entry block would freeze it at its pre-first-sample value, and there was nowhere else to put it. A benchmark bank of 16 modes sharing oneexp-derived drive term was evaluating thatexponce per mode instead of once per sample. No block is inserted: construction is CFG-linear, so every loop is already preceded by its preheader (the function prologue for the sample loop, the induction initialiser's block for afor), which leaves the block layout the wasm backend recovers regions from untouched. The single-definition guard on the hoisted instruction's own result is kept - that is what pins a promoted state register, deliberately written by both the prologue load and the per-sample move, inside the sample loop. -
YDSP optimizer: new
storeToLoadForwardingpass. A state-array read that follows a write to the same element becomes a move from the value just stored, soa[i] = x; ... = a[i];no longer round-trips through memory - in a loop over parallel arrays (the ElectricPiano partial bank readsoscI[i]back immediately after writing it) that also removes a store-to-load-forwarding stall from the accumulation chain. The rewrite replaces one definition with an identical value, so it needs no SSA property. It forwards only within a block, and only when the region, the index value and the stored value are all provably unchanged in between; an intervening array store blocks it unless that store writes a different region through the same index and element width, where the two addresses differ by their region bases alone and so cannot alias for any index. The@delay ring is unaffected: it writes and reads through different indices. -
YDSP codegen: a comparison whose only consumer is a
selectnow feeds it through the condition flags instead of a 0/1 general-purpose register.select (a > b, x, y)was emittingfcmp,cset,cmp,fcselon AArch64 (andcomiss,setcc,movzx,test,cmovon x86-64) where two instructions suffice - and thecset/cmpround-trip added two links to a dependency chain that, in a per-sample recurrence such as an envelope follower, is the critical path. The comparison is re-emitted next to the select rather than moved, so nothing can clobber the flags in between; the fusion is skipped unless the comparison is used only by that select, both are in the same block, and neither of the comparison's operands is rewritten between them. Candidates are identified by instruction position rather than by result value id, so a select written by if-conversion - whose result is a mutable local, and therefore defined more than once - fuses like any other. -
YDSP codegen (AArch64): state-array accesses no longer rebuild their region's base address on every access. AArch64 has no base + index + offset addressing mode, so reaching
array[i]means materialisingstateArrays + regionBasein a register first - which was emitted inline at each access as amovplus anadd(two extra instructions and two extra virtual registers), even though the region base is a compile-time constant. One register per distinct region is now computed once in the prologue. In a loop over parallel state arrays this was the bulk of the emitted code: the benchmark's 32-partial harmonic bank performs eight array accesses per partial, so it was emitting roughly twenty-four instructions per partial where eight are needed. x86-64 folds the same address into one addressing mode and is unaffected. -
YDSP codegen: the generated sample loop no longer pays a function call, a redundant memory round-trip or an unpredictable branch per sample. The
@delay wrap lowers to a newwrapIopcode (compare +cmov/csel/select) instead ofmodI, which on x86-64 was a call into a helper for every delay tap on every sample; scalarstateand the hidden'/@/smoothslots are loaded once in the kernel prologue and written back once in the epilogue instead of round-tripping through memory every sample; stream channel pointers are hoisted out of the loop rather than re-derived at every access; float constants are read from AsmJit's constant pool (andnegF/absFfrom a sign mask there) instead of being materialised through a general-purpose register, which also frees the registers LICM used to pin across the whole kernel;selectis branchless on both targets (fcsel/cselon AArch64,cmovplus an and/andn/or blend on x86-64); AArch64 comparisons usefcmp/cmp+csetinstead of a five-instruction branch diamond; and a terminator whose target is the next block falls through instead of emitting a jump.-0.0is the only observable behaviour change:negFnow flips the sign bit, so-(+0.0)is-0.0on x86-64 as it already was on AArch64. -
YDSP tests: new
YdspBenchmarkTeststime six discriminating patch shapes (@delay taps, a four-pole ladder, a 32-partial harmonic bank, compare +select, a modal bank with work invariant across its inner loop, and a wavefolder with data-dependentif/else) against hand-written C++ equivalents and print a ns/sample ratio for each, with a loose regression guard rather than a threshold assert. Two further shapes compare the JIT against itself to size a specific opportunity rather than to defend a number: the modal bank with and without its accumulation, whose delta is what a widened reduction costs (four serially dependent vector adds plus the horizontal fold), and three chained sample-mode nodes against one processor computing the same thing, whose gap is the headroom a kernel-fusion pass could recover - the win fusion was assumed to have, never measured. The reduction shape has since produced a more useful result than the one it was built for: removing the accumulation makes the kernel reproducibly slower (14.4 vs 20.8 ns/sample, eight fewer instructions, both loops confirmed widened), and a kernel whose runtime moves inversely with its instruction count is not issue-bound - which rules out loop overhead, and with it a full-unroll pass, as the modal bank's remaining gap. It is also the only shape in the suite with a ~30% run-to-run spread rather than under 3%, and the only one calling a transcendental inside the sample loop, so the shape now prints stack-traffic and call counts from the generated listing alongside the instruction count. Separately, the wave folder shape now also times a reference compiled with floating-point contraction disabled:last = last * 0.5f + y * 0.5fis onefmaddon AArch64 by default, which is one rounding where the source asks for two - a precision liberty YDSP does not take, because a patch must produce identical samples on every backend - so the ordinaryc++row is not running the arithmetic the JIT runs, and on a shape whose critical path is exactly one multiply-add per sample that is most of the ratio: 2.26x against the ordinary reference, 1.29x against one that also cannot fuse. The two references produce bit-identical output, because* 0.5fis exact and fusing its rounding changes nothing numerically - which is the check that the pragma took effect rather than being ignored. -
YDSP optimizer: new full-unroll pass for constant-trip-count loops. A
forloop whose trip count is known at compile time is written out in full into its preheader, which deletes the bound compare and the back edge from every iteration. It runs after the vectoriser and so unrolls at the widened trip count - a 16-mode bank at four lanes becomes four copies of a four-lane body, not sixteen of a scalar one - and running it before would have left nothing loop-shaped to widen. No block is added, removed or reordered: the header and body are emptied and left falling through, the same trickifConversionuses, so every block index, every loop bound and the region layout the wasm backend recovers are untouched; the loop's entry infn.loopssurvives so the report still answers "how many iterations could this run" with the worst case rather than 1, markedunrolledso no backend wraps a region around blocks that no longer branch. The body is copied verbatim, increment and all, rather than substituting a constant index per copy: that makes the unrolled code the identical instruction sequence the loop executed, so it is bit-exact by construction and needs no reasoning about a non-SSA IR. It declines a one-shot init kernel, a nested or runtime-bound loop, and any loop whose copies would exceed 256 instructions or 32 trips (256 admits a 32-partial bank at ~190 instructions, but that shape then moved by ~2%, inside its own spread, so the step up from 128 is unproven), and like the vectoriser it is off for wasm - a module is downloaded and parsed before it runs, and the browser's own engine re-optimises the loop regardless, so trading module size for branches is the wrong way round there. This also surfaced a latent trap in the AArch64 backend:vectorIndexRegsmemoizesindex << 2keyed by value id and cleared per block, which is only sound while a value id is written at most once per block - an unrolled loop writes its induction variable once per copy, and the stale scaled register would have addressed the wrong element with no diagnostic. A newonValueRedefinedhook invalidates the memo at every write. Measured on the modal bank, whose 16-mode loop widens to four lanes and unrolls to four copies: 2.17x slower than the C++ reference down to 1.21x (14.3 to 8.3 ns/sample). The larger result is what it did to the variance - that shape had been the only one in the suite swinging ~30% run to run, wide enough that it had reported a kernel with eight fewer instructions as the slower one and sent two investigations (spill traffic around the per-sampleexpcall, then the kernel not being issue-bound) after explanations that were not there. Unrolled, the spread is under 3%, and the shape finally produces the number it was written for: the widened reduction costs ~1.9 ns/sample, about a fifth of the kernel.YdspKernelReportgainedunrolled, sinceboundedIterationCountdeliberately reports the same worst case either way and is therefore no help in telling the two forms apart. -
YDSP optimizer: an unrolled bank's widened accumulator is now turned into a reduction tree. Unrolling leaves
acc = acc + xonce per copy and each add waits on the one before it; sending the odd copies to a second accumulator and adding the two at the end halves that serial depth for the cost of one instruction (the first odd link becomes a move rather than an add, since the second accumulator has no zero to start from, so the add count is unchanged and the combine puts one back). Halving can repeat: the scan resumes just inside the chain it rewrote and picks up the suffix still on the accumulator, so an eight-link chain halves twice and ends at depth four. That is short of what restarting the scan would reach, which is untried. It applies only to a widened accumulator: the vectoriser has already re-associated that sum - lane j sums elements j, j+4, j+8 … and the lanes are folded pairwise - and that is documented as not bit-exact, so this stays inside a licence the language already takes, while a scalar accumulator carries none and is left exactly as written. Measured onModalBankReductionCost, which times the same bank with and without its accumulation: the reduction went from 1.93 ns/sample unsplit to 0.98 and 1.53 across two split runs, and the modal bank overall from ~1.24x to 1.11-1.20x of its C++ reference. Ranges rather than figures: the direction is unambiguous and the size is not, which is also why the switch exists. That was not the expected result - the reduction looked closer to throughput-bound than latency-bound (six extra instructions, ~6.4 cycles), which is whysetReductionSplittingEnabledis a switch separate from the unroller it depends on. The repeated halving was not designed either: it was found by reading an instruction count that had grown by two combines rather than one, and kept because the largest bank in the benchmark posted its best figure with it.YdspKernelReportgainedreductionSplit, so a patch can be asked whether its reduction was shortened without inferring it from a timing.
-
GUI:
Component::hitTest()is consulted again when routing mouse events.Component::findComponentAtForMouseEvent(), which replaced the windowing layer's own traversal, tested only visibility and bounds, so a component that carved out part of its rectangle - a round knob keeping its corners transparent, a control with a padded margin - received moves, enters, exits, clicks and wheel events over the area it rejects, and whatever sits behind it received none. -
Thirdparty (
sleef_library): the SIMD translation units referenceSleef_x86CpuIDfrom their exported dispatch queries (Sleef_getIntd2/Sleef_getIntf4, theDORENAME-renamedxgetInt/xgetIntf), but the definition in SLEEF'ssrc/common/common.cwas never compiled into the module, so an x86-64 build failed to link (LNK2019forSleef_x86CpuIDinsleef_library_simddp.obj/sleef_library_simdsp.objon the Windows CI; ARM targets were unaffected because their helper header has no CPUID probe). The module now compilesupstream/src/common/common.cas its own translation unit (sleef_library_common.c), as SLEEF's own build does. -
YDSP: the four-lane x86 float comparison was miscompiled.
vectorFloatCompareemitted legacycmppswith three operands and five-bit AVX predicates, butCMPPSis two-operand (dst = cmp (dst, src)) and decodes only three bits - so the second source was dropped and the destination was read without being seeded. The path now seeds its destination and uses the legacy predicates, emittingb < afora > bsince the three-bit set has no greater-than form. Reachable on a pre-AVX2 host or undertargetPolicy = baselinewithbaselineTarget = sse2; no asmjit validation is enabled in this module, so nothing rejected it atfinalize(). -
YDSP:
INT_MIN / -1now yields 0 on AArch64, as it already did on x86-64 and wasm. The pair has no representable quotient; the x86 helpers had always guarded it, the wasm lowering guards it explicitly (div_swould otherwise trap), and the AArch64 lowering - whose comment claimed parity with x86 - guarded only the zero divisor and letSDIVwrap the quotient toINT_MIN. The same patch therefore produced a different value on Apple silicon than on an x86 host.CMN b, #1gates theINT_MINcompare so the common path pays one extra compare, on an operation the@delay wrap no longer uses at all. -
YDSP: integer division and modulo on x86-64 no longer make an out-of-line helper call per occurrence.
emitIntDivisionnow inlinescdq/cqoplusidivbehind the same zero-divisor guard the AArch64 lowering uses, so both targets keep returning 0 for a zero divisor, plus a guard forINT_MIN / -1- the one pair with no representable quotient, which raises#DEon x86 rather than wrapping. The orphaneddivInvokemember and the fouryupDspIdiv/yupDspImodhelpers go with it. -
YDSP optimizer:
splitWidenedReductionChainsnow snapshots each link's addend to a fresh value at the link itself before building the balanced tree. Loop unrolling duplicates a body verbatim, so one value id is redefined once per copy; the tree is emitted after the whole chain, so every leaf previously resolved to the last copy's definition and was summed repeatedly - the ElectricPiano harmonic bank summed only its final 4-lane product, making the patch sound as a thin ~0.1 s high-frequency click under the automatic tier (scalar and baseline tiers were correct, which is why only the widened demo patch was affected). The newYdspElectricPianoTests::SustainsPastTheAttackClickand theYdspExamplePatchTestsElectricPiano sustain tests are the regression guard. -
YDSP analysis: each top-level body (
process,init) is now its own local scope, so a same-named local in both no longer reports a spurious duplicate-symbol error; a node writtenP * 0/P / 0is rejected with a diagnostic instead of reaching the latency math with a zero rate factor. -
YDSP optimizer:
if (c) { y = a; } else { y = b; }where each arm moves into the same value is now fused into oneselect(the wavefolder / clipper shape), removing the branch.pow (x, 2.0)becomesx * xunder fast math (strict keeps the intrinsic, since libm pow is not bit-equal to a multiply). Graph meter reads use compile-time byte offsets instead of re-scanning the meter type list per getter. A spelledfma()/fmsub()no longer disqualifies an otherwise widenable loop on targets with a packed fused multiply-add (x86 AVX2+FMA and AArch64), where the widened op lowers tovfmadd213ps/fmla. -
YDSP runtime: output events carried across a block boundary are stored relative to the next block's start and delivered directly, so an event that straddles into a differently-sized block (64 then 32 samples) still lands at its exact sample instead of being unwound against the wrong block size. The host-side automation and MIDI sample offsets are also clamped to
[0, blockSize - 1]rather than the inclusiveblockSize. -
YDSP runtime: a polyphonic node's shared per-voice scratch is cleared for each awake voice before it renders (previously once per block), so a voice whose kernel reads its own output (
out = out + 1.0) or writes it conditionally no longer sees the previous voice's samples - two held feedback notes now sum to exactly 2.0 per sample instead of leaking 3.0. -
YDSP runtime:
process()no longer silently truncates a request larger than the prepared block size; it returns the newYdspProcessResult::blockTooLargeand leaves the graph untouched.reset()andprepare()now clear the dropped-event counters (the host-visibledroppedEventCountand every node's output-event queue count), which previously grew monotonically for the life of the graph. -
YDSP runtime:
accumulateStreamrefuses (rather than reinterpreting as float32) when a non-float stream ever reaches the summing fan-in path; the analyzer already rejects such graphs. -
YDSP wasm backend: integer
div/modnow also guardsINT_MIN / -1(wasmdiv_s/rem_straps on it, like a zero divisor), mirroring the x86/AArch64 helpers. The emscripten kernel registrar checks its registry before compiling, soprewarmKernels()no longer recompiles an already-registered module. -
YDSP runtime:
setParameter/setDoubleParameter/setIntParameterno longer memcpy straight into the params block the audio thread reads; they push (slot, raw bits) into a fixed control-to-audio ring thatprocess()drains once at the top of the block, so host parameter writes are race-free against the audio thread and land at the start of the next block (an overfull ring counts as a dropped event). The typed getters drain the ring first, so a set is immediately visible to a control-thread get. -
YDSP codegen: the operand-register diagnostic skips scalar-load slot operands (operand 0 of
loadParam/loadParamOut/loadStateF/loadStateIholds a state/param slot, not a value id), so state-heavy kernels no longer false-positive as having an undefined value. -
YDSP optimizer:
foldStateWriteBacksno longer folds a scalar state write-back whose produced value is read outside the move's own block - a function parameter reassigned inside anifkeeps being read by the join block after the call, and folding the parameter copy renamed its producer and left that read with no definition at all, which asmjit rendered as<Reg-0>?255on AArch64 (the tenYdspExamplePatchTeststhat failed that way are the end-to-end guard). The fold also stops treating a scalar load's slot index as a value id, requires matching value shapes, and applies overlapping write-backs into the same state register one at a time instead of as a batch against a stale snapshot. Dead-code elimination and the contraction pass now share one filtered value-use scan (forEachValueUse), so slot indices can no longer be mistaken for value uses anywhere. -
YDSP optimizer: confirmed miscompiles fixed. Loop fusion retargets the fused loop onto the exit block its own loop record names (it used to aim the header one block before the loop, at the preheader), only drops a second loop's induction init when it really starts at the literal 0 (a
3..8second loop no longer silently restarts at 0), and treats event emission and param/paramOut/event-field accesses as memory - two event-emitting loops no longer fuse. Constant folding no longer materialises a fold its own guard declined (x / 0,x % 0,x << 64stay unfolded instead of becoming a compile-time 0), refuses to fold a non-finite or out-of-rangeftoi, and rounds folded float32 values to float32 so folded and unfolded kernels agree. Strict math no longer applies thex * 0.0 -> 0.0/x + 0.0 -> xidentities (both are wrong for signed zero, NaN and infinity). Copy propagation no longer retargets a block's terminator condition past a redefinition of the moved register. A widenedfmaFis refused by the scalar lowering instead of being scalarised per lane, a widened-reduction link must be single-use to be deleted, and dead event-field loads are now removed by dead-code elimination. -
YDSP native codegen: the temporary
YUP_YDSP_RA_DEBUGenvironment switch was removed (RA annotation is compiled out), a scalar-transcendental table miss now asserts instead of silently leaving the result register undefined, and the post-register-allocation redundant-move cleanup also removes identicalfmov/ASIMDmovcopies. -
YDSP codegen (x86-64): integer variable-count shifts wrote the fixed physical
clregister directly (mov cl, srcB.r8()). AsmJit's register allocator is blind to physical operands - it only models virtual ones - so it could keep another live value in RCX and themov clsilently destroyed its low byte (most often a LICM-hoisted shift-count constant), corrupting the LFSR generator'sw >> k/<< 31chain. The count is now passed as a virtual register and the allocator pins it into CL with full liveness knowledge. -
YDSP codegen (x86-64): the shared
clampIlowering selected the max into the result register and then ran the min with that same register as its ownwhenTrueoperand. The x86-64 select lowering writesdstbefore readingwhenTrue(mov + cmov), so the second stage degenerated to a no-op andclamp (n, -10, 10)always returned the upper bound -min/max/clamp/abs/signon an int returned6where-14was expected for negativen. The max stage now selects into a temporary, keeping the two select operands distinct on both backends. -
Emscripten:
PopupMenu's click-outside-to-dismiss and hover-forwarding (Desktop::handleGlobalMouseDown/Up/Move, fed from an SDL event watch inyup_Initialisation_sdl.cpp) built theirMouseEventfromSDL_GetGlobalMouseState(), which reports real page-relative pixels on this platform.Component::localToScreen()instead anchors onSDLComponentNative::getPosition(), which returns(0,0)on Emscripten by design (there's no real window position to query, see theisMouseOutsideWindowfix below). The two frames only agreed when the canvas happened to sit near the page origin, so selecting aComboBox/PopupMenuitem would misfire "click is outside the popup" and dismiss it before the selection could register - reproducible on both HiDPI and standard displays, and sensitive to anything that shifts the canvas's on-page position (such as opening devtools). The global dispatcher now sources its position the same way regular per-window dispatch does on Emscripten, keeping desktop behavior (which needs true desktop-global coordinates) unchanged; mobile is left untouched pending separate verification. -
Emscripten:
isMouseOutsideWindow()comparedSDL_GetGlobalMouseState()(real page-relative coordinates) againstSDL_GetWindowPosition(), which is meaningless on this platform sinceSetWindowPositionis unimplemented in the SDL3 Emscripten backend and returns a value centered against a synthetic virtual display. This misclassified ordinary clicks inside the canvas as "outside the window" on every mouse-up, which calledhandleFocusChanged (false)and stopped SDL text input - silently dropping every typed character while raw key events (Tab, Backspace, Enter, and anything readingKeyPressdirectly, likeMidiKeyboardComponent) kept working, the exact symptom reported against the newCodeEditor. The check is now skipped on Emscripten, where "mouse outside the window" isn't a coherent concept for a single always-present canvas. -
YDSP codegen (AArch64): the register allocator was free to park a long-lived value in x30 (LR), which any call then destroyed. asmjit's AArch64 tables disagree with AAPCS64 about that register:
a64func.cpplists x30 among the callee-preserved GP registers, so the allocator believes a value there survives a call, whilea64rapass.cpponly makes SP and FP unavailable - leaving x30 allocatable.blroverwrites LR with the return address, so any kernel that both calls out (a transcendental, or theexpbehindsmooth) and has enough simultaneously-live values to reach x30 had that value silently destroyed mid-body. The symptom depended entirely on what the lost value was doing: a state-array region base meant every later access through it addressed wild memory, which is how this was found - a fused kernel with four reverb-sized@rings and a smoothed parameter faulting onstr s26, [x30, x15, lsl #2]- but a promotedstatescalar in x30 faults on nothing and simply makes the patch compute the wrong numbers, so this was a latent wrong-audio bug of unknown age rather than only a crash. x30 is now added to the function frame's unavailable set, which costs one of ~29 allocatable GP registers and is what a JIT should do with LR regardless; the fix sits on our side of the boundary so it survives an asmjit update. It needed three ingredients at once (a call, high register pressure, and a state array large enough for the corruption to land outside the allocation), which is why every existing patch and test missed it: four newYdspFusionTestsshapes bisect exactly those ingredients, and the crashing combination is now a regression test. -
YDSP runtime: three routing defects that only the old exactly-once rule kept unreachable, all fixed by the same new representation (a node's input wiring and the graph's output wiring are both CSR-flattened lists of connections, each owning its own delay ring). (a)
outputSlotBuffer[s]was last-edge-wins, so a node output feeding both a graph output and a node input clobbered one marker against the other - and the two edge orderings failed differently. (b)outputSlotGraphOut[s]was a singleint, so one output could not drive two graph outputs at all: the second silently replaced the first. (c)edge.delaySampleswas read only on the node-input branch, soa.out -> [3] -> ywas parsed, accepted and then silently dropped. Owning the delay ring per connection rather than per input slot is what makesx -> a.in; x -> [4] -> a.in;a comb filter instead of representable-only-once. -
YDSP runtime: a
connection { x -> y; }passthrough edge (graph input straight to graph output, whether hand-written or produced by inlining a pass-through subgraph) hit neither branch of the wiring producer, so it compiled and then output silence. It is now an ordinary graph-output source with no producing node. -
YDSP: oversampling (
node x = P * 4) on a non-float32stream was silent memory corruption.rateMultiplierwas validated only for the event-driven and subgraph cases - there was no type guard anywhere - and the runtime handed the node's buffers toOversampler<float, ...>, so afloat64stream was read and written at half stride, past the end of the buffer. It is now a compile error naming the offending stream and its type. Inline delays already carried the equivalent guard. -
YDSP: two out-of-bounds reads in the graph algebra's sequential composition. A graph input or output identifier used as a leaf gets arity
(1, 1)but carries a port on one side only, sodry : wetproduces a value claiming arity(1, 1)with no ports at all - the arity check then passes fordry : wet : Gainand the wire loop indexedoutPorts[0]on an empty vector. Its mirror image,Gain : dry, indexedinPorts[0]. Separately,_inside,contributes no ports while still adding to the arity, so(_ , Gain)- and(Gain , _)- claimed arity 2 with one port, and whatever was sequenced against it readinPorts[1]. The composition now verifies that a declared arity is backed by actual ports and reports which side has nothing to connect, and mixing_with a port-carrying operand inside,is rejected with a message pointing at theconnectionform.(_ , _)still composes as it always did: it stays an identity. -
YDSP runtime:
prepare()appended to each node's inline-delay buffer list instead of rebuilding it, while the per-block path indexes that list 1:1 with the input slots. A secondprepare()therefore left every slot pointing at the first call's ring, whose one-block scratch is sized for the old block size - so preparing at a larger block size wrote past the end of it. The list is now assigned to the input-slot count and filled by index, which removes the mismatch by construction.reset()also left the rings holding audio (the oversamplers beside them were already reset), so the pre-reset tail kept playing out fordelaySamplesmore samples; they are cleared now. Both were uncovered because every existing test prepares exactly once. Relatedly,prepare()ran the one-shot init kernels before rebuilding the scratch arena and the node pointer tables; the init kernels are handed null stream pointers so nothing observable changed, but they now run last. -
YDSP optimizer:
algebraicSimplificationtreated any constant writing a register as that register's value, ignoring that the IR is not SSA and the register may be written again before the use being simplified.process { float t = 0.0; t = in * 2.0; out = t + 1.0; }foldedt + 1.0againstt's declaration literal and emittedout = 1.0, dropping the input entirely - silent wrong audio, no diagnostic. Only a register defined exactly once is eligible now, which is the guardconstantFoldingalready documents and depends on. The bug needed the reassignment and the later use to share a basic block, which is why the commonfloat sum = 0.0;followed by an accumulating loop was never affected (a loop body is a separate block) - the kernel-fusion pass, whose junction locals are declared, assigned and read in one block, is what exposed it. -
YDSP runtime: a connection with an inline delay (
x -> [3] -> node.in) wrote the delayed samples back over its source buffer. For an edge fed by a graph input that is the caller's own array, whichprocess()declares asSpan<const float>and the runtime wasconst_cast-ing away: a host that reuses or shares an input buffer found it silently rewritten after every block, and a genuinely read-only buffer was undefined behaviour. The delayed samples now go into a per-edge scratch buffer allocated inprepare()and the node's input pointer is repointed at it, so the source is untouched; the sub-block dispatch paths re-derive from that pointer and needed no change. The feature had no test anywhere in the suite, which is why this survived - it now has three, including one asserting the input buffer is byte-identical afterprocess(). The delay ring is float32-only (as oversampling already was), so an inline delay on a non-float32 stream is now a compile error instead of reading the buffer at the wrong element width. -
YDSP codegen (x86-64): the backend did not compile at all. It told a float32 register apart from a float64 one with
reg.typeId(), which the vendored AsmJit has no such member for - the API is snake_case throughout (is_vec128,new_gp32,virt_reg_by_reg) - so all thirteen call sites were hard errors, hidden on AArch64 hosts by the#if ASMJIT_ARCH_X86guard around the file. The width was never available on the operand in the first place: unlike AArch64, x86 has no distinct register class per float width, so a f32 scalar, a f64 scalar and a packed f32 vector are all a 128-bit Xmm. It is read from the register'sVirtReginstead, as theTypeIdit was created with. -
YDSP runtime: mid-block parameter automation on a polyphonic node reached voices 1..N-1 an entire sub-block early. A node's param block is shared by all its voices, but
applyAutomationruns inside the per-voice loop and nothing rewound the block between voices, so voice 0 saw the correct pre/post values while every later voice entered the block with the param already holding voice 0's post-automation value - and therefore read the new value throughout its pre-automation sub-blocks. AVoice[4]bank with two notes held and a gain automated from 0.5 to 1.0 at sample 100 produced 1.5 before the offset instead of 1.0. The block-start value of every automated param is now snapshotted before the voice loop and restored at the top of each subsequent voice, so all voices replay the same timeline. Monophonic and single-voice nodes run the loop once and are untouched. -
YDSP optimizer: a loop nested inside another loop overwrote the outer loop's recorded
YdspLoopBound. The bound is resolved before the body is lowered but only stored on theYdspIrLoopafterwards, and it was held in a builder member in between, so both loops ended up reporting the inner bound andYdspKernelReport::boundedIterationCountunder-reported the worst-case iteration count. The bound is now captured into a local at the point it is resolved. -
YDSP runtime: a node's all-sound-off handling (MIDI CC120) kept a single sample offset per node per block, so two CC120s on different event inputs clobbered each other. The earlier (more urgent) silence point was lost, and a note re-triggered between two CC120s escaped the later one. Every CC120's offset is now recorded (fixed per-block capacity, reserved in
prepare()) and applied as its own sub-block split point. Voice lookup is also scoped by event input now: each event input has its own MPEInstrument, so the same MIDI channel+note collides across inputs, and a note-off on one input could previously release a voice (poly) or remove a held note (mono) belonging to the other input. -
YDSP compiler: syntax errors in an imported file were swallowed. The parser recovers from a syntax error via
synchronize(), so the merged program was non-null and the import merged anyway, dropping the diagnostic; the merge now treats any recorded error like the top-level path does. The import rename pass now also covers processor-local function bodies, graph node parameter overrides and graph-leaf override expressions (plain-name library calls in those positions previously dangled after merging). The static recursion check now walks if/loop bodies and calls nested anywhere in an expression, so a recursive library function hidden in a conditional is rejected before inlining. -
YDSP compiler: a diamond import silently dropped the second alias's namespace. The merge deduplicated by file path globally, so when two imported files both imported the same third file, only the first importer's namespace was populated - the second importer's references to its own alias of that file failed to resolve (and a same-file-different-alias import in one file lost the second alias too). The merge now deduplicates per (importing file, file, namespace) and runs each merge on a private copy of the parsed file, so every importer gets the shared declarations under its own namespace.
-
YDSP compiler: the intrinsic name list was hardcoded twice - a mirror of the semantic analyzer's table lived in the compiler's import rename pass, so the two could drift (and already had, once). The rename pass now queries the analyzer's table through a shared
isIntrinsicName(). The function shadowing rule (a processor-local function beats a program-level one of the same name) was likewise duplicated between the semantic analyzer and the IR builder; both now resolve through one sharedfindFunctionInScope(). The module's runtime-initialised static lookup tables (intrinsics, builtin constants, IR op dispatch, annotation keys) are nowconstexprarrays with no dynamic initialisation. -
yup_core: new
CountDownLatchsynchronization primitive (addCount/countDown/wait), used by the parallel import parser to replace its hand-rolled mutex/condition-variable counting; the threaded import path now also has test coverage for diamond imports and failure cases. -
YDSP runtime: sample-accurate sub-block dispatch advanced each sub-block's stream pointers from the previous sub-block's pointers while passing an absolute offset, so the offsets accumulated: any node split into three or more sub-blocks in one block (two or more distinct event/automation offsets) wrote its output - and read its input - at
sum(offsets)instead ofoffset, corrupting samples past the end of the buffer in the worst case. Polyphonic outputs were immune (they re-derive from the per-voice scratch each call), which is why the existing voice-bank tests never caught it. Each sub-block now re-derives its pointers from the block's base. -
yup_gui: mouse-event hit-testing now respectssetWantsMouseEvents()across siblings, not just up the parent chain. NewComponent::findComponentAtForMouseEvent()walks the hierarchy the same wayfindComponentAt()does, but a component (or its whole subtree) that opted out of mouse events is skipped in favor of the next sibling underneath it, instead of only bubbling up to a parent;findComponentAt()itself is unchanged (still a pure, mouse-event-agnostic geometric hit-test). Previously an overlapping sibling that didn't want mouse events (e.g. a decorativeLabeldrawn on top of aSlider) would still win hit-testing and swallow clicks meant for the sibling beneath it.SDLComponentNative::findComponentForMouseEvent()now delegates tofindComponentAtForMouseEvent()instead of duplicating the ancestor-walk itself. -
yup_core:Thread::getCurrentThreadId()is now TLS-based on wasm instead of returning the rawpthread_self()value, which is unreliable on the emscripten audio-worklet thread (a Wasm Worker, not a pthread) - thread-identity-based locks and checks (RecursiveSpinLockowners,ReadWriteLockwriters,MessageManagerthread checks) misbehaved there. Every thread (main thread, audio worklet, yup::Threads) gets a stable TLS id seeded by an atomic counter; yup::Threads sync theirthreadIdto it at creation sogetThreadId()keeps matching, whilethreadHandlestays the native pthread handle. The implementation must not consultgetCurrentThread()-ThreadLocalValueusesgetCurrentThreadId()to find its per-thread slot, which would recurse. -
yup_audio_basics:AudioLockType(used by theMidiKeyboardState,SynthesiserandBufferingAudioSourcelocks) is aRecursiveSpinLockon wasm (never blocks - nofutex_wait- and re-entrant, whichSynthesiserrequires when processing MIDI) and the re-entrantCriticalSectionon desktop.MidiKeyboardState::allNotesOffadditionally no longer re-enters the lock (allNotesOff/noteOffdelegate to non-recursive*Lockedhelpers), so the state is safe even with a non-recursive lock. -
yup_audio_devices: the locks taken on the audio thread -AudioDeviceManager::audioCallbackLock/midiCallbackLock(taken in everyaudioDeviceIOCallbackIntblock),MidiMessageCollector::midiCallbackLock,AudioSourcePlayer::readLockandAudioTransportSource::callbackLock- now useAudioLockTypeinstead of a rawCriticalSection, so on wasm a contended lock spins on the audio-worklet thread instead of blocking in a fatalfutex_wait.getAudioCallbackLock()/getMidiCallbackLock()now returnAudioLockType&. -
YDSP wasm codegen: stream loads/stores (
loadInput/loadOutput/storeOutput) treated the channel-pointer array (inputs/outputs) as the sample buffer - the channel pointerinputs[ch]was never dereferenced, so a kernel indexed&inputs[0] + i*4and wrote sample data over the pointer array and into the surrounding heap. Every generated kernel that touched a stream corrupted the heap (crashes surfaced in the browser's audio worklet and in node tests). The channel pointer is now loaded first, then indexed by the sample. -
examples/graphics: the "Pulse Bass" patch is now the demo's monophonic example -node voices = PulseVoice [[ mode: mono, priority: last ]]- and is the first patch to reade.isLegato: a fresh key press re-attacks at the new pitch immediately, while a legato continuation leaves the envelope running and glides the pitch over a newGlideparameter. The re-attack ramps from wherever the envelope currently is rather than resetting it to zero - with a single voice, resetting would chop a still-sounding release tail off mid-level and click, and ramping is how a mono synth's single-trigger envelope behaves anyway. Previously it was a 4-voice bank whose.isLegatoflag nothing consumed, so the mono/legato path shipped untested by any example. -
examples/graphics: the "YDSP Synths" patches clicked on every note.Poly Sine,Analog Saw,FM BellandWobble Leadall stepped their envelope straight to the note velocity innoteOnwhile the oscillator kept its previous phase, so the output jumped discontinuously at each attack (and at each voice steal). They now run a ~3 ms one-pole on the gain - the coefficient is astate float ... = 1.0 - exp (-samplePeriod / 0.003)initialiser, so it is computed once per voice in the init kernel rather than per sample.Pulse Bassalready had a one-pole attack and is unchanged. -
examples/graphics: the "Electric Piano" patch was ~24 dB quieter than the other synths (a* 0.125output trim on top of partial weights that already sum to ~1); it now matches their level. -
examples/graphics: the "YDSP Synths" demo only played patches with exactly one output stream - any other count fell into the silence branch, so a stereo patch compiled and ran but was never invoked at all (process()also rejects a buffer span whose size does not match the declared stream count). It now accepts mono and stereo patches: a mono graph is fanned out to every device channel as before, a stereo graph maps its two output streams to alternating channels, and the oscilloscope shows the mean of the two. -
YDSP optimizer: a
forloop nested inside anifhung the generated kernel. The if-lowering allocated its join block (and, with anelse, the else block) before lowering the then region, so any block the region allocated landed past the join: the loop's preheader fell through into the join instead of into the loop header, and the join then fell through into that header - whose induction variable had never been initialised - so the loop ran forever. Both the join and the else block are now allocated after the regions they follow, keeping each region's blocks contiguous and the join last, which is the layout both the fallthrough-based asmjit codegen and the wasm backend's structured-region recovery require. Theelse iffix below only covered the else-carrying case. -
YDSP optimizer:
if / else ifchains produced a non-linear block order - the outer join block was allocated before the inner if's blocks, so the outer join fell through into the inner then-block and the else-tail branch looped back into a cycle (infinite loop on the asmjit backend; "unexpected conditional branch" compile error on the wasm backend, which surfaced it first via the AnalogSaw patch'spolyBlep). The join is now allocated after the else region, so nestedelse ifblocks stay between the else block and the join, restoring the CFG-linear order both codegens rely on. -
YDSP optimizer:
loopInvariantCodeMotiontreated loop-carried registers as invariant - the induction register is written both before the loop (sample mode: prologueconstI 0; block mode: preheadermovI) and inside the loop body, so itsaddI/movIupdate was hoisted out of the loop, freezing the counter and hanging the generated kernel (the sample-mode case slipped through even after the preheader exclusion because the prologueconstI 0is a constant op that re-seeded it). Only single-assignment values defined in the entry block (or constants) can now be invariant, registers redefined inside the loop (induction variables, path-dependent if/else values) are never hoisted, and loads from input/output streams, params and meters are not hoisted either (the loop body can overwrite that memory). -
YDSP optimizer:
constantFoldingtreated the non-SSA IR's value ids as single-assignment, so the sample-loop induction register - defined once by the prologueconstI 0and once per iteration by themovIupdate - folded its per-iteration update to a literal and the generated kernel looped forever (CompilesPassThroughAndProducesCorrectOutputhung). Constants are now only propagated for value ids defined exactly once. -
YDSP optimizer: inlined
funccalls bound parameters directly to the caller's argument register (locals are mutable single-register IR slots), so a function that reassigns a parameter - e.g.t = t / dtin the AnalogSawpolyBlep- clobbered the caller's local and the block-end state store persisted the mutated value instead of the original (state[0] = tinstead ofphase), making the oscillator diverge to ±inf after the first cycle. Parameters are now pass-by-value: each argument is copied into a fresh value at the call site. Copy propagation also stops at redefinitions of the destination, so it can't propagate through the new parameter copies. -
MidiKeyboardComponent: computer-keyboard input now actually matches keys (the previous full-KeyPresscomparison never matched SDL events) and sends matching note-offs onkeyUp, releasing all keyboard-triggered notes when focus is lost instead of leaving them stuck; the mouse wheel now scrolls left/right by single white keys (vertical wheel works on horizontal keyboards too) and Ctrl/Cmd + wheel zooms in/out, clamped between a single octave and the full 0-127 range. -
YDSP optimizer: binary ops mixing a non-literal constant with a literal (e.g.
int64(1) << 32) previously narrowed the constant to the literal's width, silently changing the result type (int64(1) << 32computed as 32-bit and wrapping to1);unifyOperandsnow mirrors the analyzer's contextual literal adaptation and only widens the source literal. -
YDSP codegen: parameter and state loads now consistently read their slot index from the
aIR field (wasmemIndex), fixing invalid memory offsets for params,prev/memslots,@delay write pointers and param-out meters. -
YDSP optimizer: implicit
int -> floatcoercion for literals and values used in float arithmetic, comparisons, intrinsics, and ternary/select branches; fixes invalid mixed integer/float instructions on AArch64 (fmul s0, s0, w0and similar).sign()now converts its integer result to a float. -
YDSP optimizer: the
@delay primitive delayed byn + 1samples; the ring read now targets the slot one past the write pointer sox @ ndelays by exactlynsamples. -
YDSP codegen:
fmod(a, b)had no codegen case and silently produced garbage output; now lowered to afmodflibm call. -
YDSP codegen: integer
/and%by a zero divisor now consistently return0on both x86-64 and AArch64 (previously trapped withSIGFPEon x86-64, and%returned the dividend unchanged on AArch64). -
YDSP codegen: libm calls (
sin,cos,tanh,pow,fmod, etc.) emitted an invalidblrwith an immediate operand on AArch64; the call target is now materialized into a register first. -
YDSP optimizer: mixing an integer literal with a float expression (e.g.
1 - aor1 + 0.5 * side) lowered to integer arithmetic, coercing the float operands throughfcvtzsand truncating fractional values; integer literals now adapt to the other operand's float width as the language's contextual literal adaptation requires. -
YDSP semantic analyzer: writing 64-bit (
float64/int64) input-value parameters in sample mode is now allowed (they double as per-sample accumulator registers); 32-bit parameters remain block-rate only. -
YDSP semantic analyzer: a processor with two or more
funcdefinitions crashed during analysis -functionTablestored pointers into the growingfunctionsvector, so registering the second function reallocated it and dangled the first function's entry (resolveFunctionCallthen read freed memory). The vector is now pre-sized before registration. -
YDSP optimizer:
copyPropagationrewrote thea/b/coperand fields of every instruction, conflating value ids with the state/param slot indices carried byloadStateF/loadStateI/loadParam/loadParamOut(both live in the same small-int space). A sample-mode state reassignment redefines a low value id through a mov, so in a Reverb-style kernel (twelve@delays interleaved with comb-state reassignments) the hidden ring write-pointer loads had their slot rewritten to a value id and read the pointer from inside the array segment instead of the scalar slot (loadstate+264vs storestate+108) - the ring store then indexed with garbage and crashed the generated kernel on the audio thread. Copy propagation now only rewrites operand fields that hold value ids for the target instruction. -
yup_core(wasm):Time::getMillisecondCounterHiRes()- read every audio block byMidiMessageCollector::removeNextBlockOfMessages- calledstd::chrono::steady_clock::now(), whichstd::terminate()s (surfacing as an opaqueabort()) when called from the browser's AudioWorkletGlobalScope: libc++'ssteady_clock::now()goes through a pthread-awareclock_gettimepath, and the audio-worklet thread is a lightweight Wasm Worker rather than a real pthread - the same "not actually a pthread" gap already worked around elsewhere for thread identity (Thread::getCurrentThreadId(), TLS-based on wasm) and for locking (AudioLockType, a non-blockingRecursiveSpinLockon wasm).getMillisecondCounterHiRes()/getHighResolutionTicks()/yup_millisecondsSinceStartup()now go throughemscripten_get_now()on Emscripten instead - a thinperformance.now()wrapper with no pthread machinery behind it, safe on every thread including the audio worklet. (An earlier attempt at this fix, changing the startup-time baseline from a function-local to a namespace-scope static, addressed a real but unrelated latent issue and did not fix this crash - the terminate was insidesteady_clock::now()itself, not in its own fallback baseline.) -
YDSP semantic analyzer:
procStates(parallel to each state's per-processor index) was never cleared between processors, so a program with two or more processors declaring state could resolve a later processor's state against an earlier processor'sYdspStateDecl- wrong type, wrong struct, silently wrong analysis. It is now reset alongside every other per-processor field. -
YDSP optimizer:
constantFoldingcomputeddivI/modI/shlI/shrI/addI/subI/mulI/negI/absIthrough plainint64_tarithmetic, which is undefined behaviour at the extremes (INT64_MIN / -1, a shift amount outside[0, 63], signed overflow nearINT64_MIN/INT64_MAX) - triggered by the compiler's own fold, not generated code.divI/modI/shlI/shrInow leave the instruction unfolded (like the existing div-by-zero guard) when the operands would invoke UB; the others compute throughuint64_tso the result matches the wraparound generated code would produce at runtime.YdspSemanticAnalyzer::tryConstantFold'sshl/shr(used when folding a program-levelletconstant) had the identical unguarded shift and is fixed the same way. -
YDSP runtime:
getActiveVoiceCount()read a voice's[[ role: voiceActivity ]]flag throughstate.data() + offsetunconditionally -stateis only sized byprepare(), so calling it right aftercompile()(isValid()is already true then) read past the end of an empty buffer. It now falls back to the same conservative "treat as active" default used for a held voice. -
YDSP runtime: an all-sound-off event (MIDI CC120) dropped for capacity (more than 16 in one block on the same node) still wiped the node's voice-slot bookkeeping unconditionally, freeing those voices for reallocation even though
silenceVoice()never ran for the dropped offset - a note reusing the slot inherited the previous voice's filter/envelope state instead of a clean one. The bookkeeping is now only cleared when the event itself was actually recorded. -
YDSP language: the recursive-descent parser, the semantic analyzer's AST walk, the IR builder's function-inlining chain and the wasm codegen's block/loop/if emission had no recursion-depth guard, so a pathologically nested expression, statement or call chain (not necessarily adversarial - a generated
.ydspsource nests easily) could overflow the native stack instead of failing with a diagnostic. All four now share oneYdspRecursionGuarddepth limit. -
YDSP language: the lexer narrowed each source character to
unsigned charinpeek()/current(), so a non-ASCII character whose low byte matched a whitespace/newline byte could end a//comment early (leaking the rest of the comment as real tokens) or desync line/column tracking; string-literal content was separately narrowed tochar, corrupting non-ASCII text. Both now carry the full code point. -
YDSP language:
parseProcess()/parseEventHandler()pushedparseStatement()'s result into the process/event-handler body unguarded, unlike the otherwise-identical pattern inparseBlockStatement()/parseFunction()- a malformed statement (e.g. a bare;) landed as a null entry in that vector.YdspParser::parseProgram()'s doc comment also claimed it returns nullptr on any syntax error, which it never does (parsing recovers and returns a non-null, structurally complete program - callers must checkdiagnostics.hasErrors()); the comment now matches the actual, intentional contract. -
YDSP runtime:
droppedEventCountwas a plainuint64_tincremented on the audio thread and read from a UI/control thread with no synchronization - now a relaxedstd::atomic. Several byte-buffer reads/writes (constant folding into state, parameter automation snapshot/restore,get/setParameter, meter reads) aliased auint8_t*through a differently-typed pointer, which is undefined behaviour when the buffer offset isn't aligned for the target type; all now go throughmemcpy, matching a precedent already used elsewhere in the same file. -
YDSP backend:
YdspCompiledKernel's wasm module lookup (wasmModules/wasmIndex) had no bounds check, so a kernel wrapper that outlived a graph rebuild could index past the end of the module vector; it's now bounds-checked before every access. Function-pointer conversions were scattered across several ad-hocreinterpret_casts (contradicting a comment claiming they were "confined" to one class) and are now centralized behind two small helpers (ydspFnPtrCast,ydspFnPtrToInt64) inyup_YdspAbi.h. The wasm text decoder's LEB128 reader could shift a 32/64-bit accumulator by more than its width on a malformed byte sequence (UB); it now caps the shift. -
YDSP backend (emscripten): a pointer was narrowed to the wasm-call context id via a direct
reinterpret_cast<int>, which isn't a well-defined conversion in general (it happens to work only because this file is wasm32-only, where pointers andintare both 32 bits) - now goes throughuintptr_tto make that dependency explicit.yupDspWasmFreeKernelalso skipped the same registry-exists guard every other function in the file uses, so freeing a kernel in a realm that had never registered/called one threw instead of no-oping. -
YDSP codegen (AArch64): scalar
float32values were allocated through the packedkFloat32x1TypeId, which AsmJit maps to a 64-bit (D) register on AArch64 but a 128-bit XMM on x86-64. Every float32 value was then computed at 64-bit width, loads over-read the next slot, and scaled array/stream stores (str d0, [base, index, lsl 2]) failed to assemble withInvalidAddressScale, breakingAnalogSawon arm64. The scalar TypeIds are used again on AArch64; x86-64 keeps the packed forms. -
yup_audio_gui:MidiKeyboardComponentrepainted synchronously from itsMidiKeyboardState::Listenercallbacks, but those callbacks are documented to be able to fire from an audio or MIDI input thread - feeding the state viaprocessNextMidiEvent()from a CoreMIDI callback repainted off the UI thread, racing the paint loop (stuck key highlights) and crashing inComponent::repaint. The listener callbacks now only trigger an internalAsyncUpdater; the repaint is coalesced and applied on the message thread, so the key state may be updated from any thread. Regression-tested inyup_MidiKeyboardComponent.cppby feeding notes from a worker thread and asserting the deferred repaint lands on the message thread. -
Examples: the YDSP Synth Lab demo (
examples/graphics/source/examples/YdspSynths.h) registered itsMidiMessageCollectordirectly as theMidiInputCallback, so hardware MIDI went straight into the audio queue and never reachedMidiKeyboardState- incoming notes never highlighted the on-screen keyboard, and incoming pitch bend / CC1 never moved the wheels. The demo is now the callback itself: it feedskeyboardState(which forwards note on/off to the collector through its existing keyboard-state listener), maps incoming pitch bend and CC1 ontoPitchWheelComponent/ModWheelComponentwithdontSendNotification(applied inrefreshDisplay()on the message thread), and forwards every other message to the collector. -
YDSP wasm backend: three defects fixed. The vectoriser widens float compare/select chains by reusing the scalar
eqF..geF/selectBopcodes onlanes > 1values, and the wasm codegen lowered every compare to a scalarf32.gt/f32.eqoverv128operands (the only branch was 32/64-bit) and a widened select to the untypedselect- so any vectorized kernel containing a comparison produced a module the browser refused to compile (expected type f32, found local.get of type v128, theYdspGraphTests/YdspVectorizerTestscompare shapes). The codegen now emits thef32x4.eq..geper-lane mask opcodes and lowers a widened select throughv128.bitselect, matching the AsmJitcmpps/fcmeqplusandps/bslsemantics. Second, wasm ignoredfastMathentirely (it was clamped false), soa * b + cstayed two separate float32 roundings -TheContractedSubtractRoundsOnceproduced0.789999962instead of the fused0.790000021. wasm now honorsfastMathlike native: scalar float32 contraction is fused there too, expanded through the exact float64 sequence (the target has no fused multiply-add instruction), so it rounds once and stays bit-stable with the native default. Because a widened chain cannot fuse without a packed fused instruction, the vectoriser keeps the implicit per-sample stream loop (runtime blockSize bound) scalar when it holds a fusable mul->add/sub chain on such a target (new rejection reasonkeptScalarForContraction) so the chain still rounds once; constant-bound bankfor i in 0..Nloops keep widening unfused by design. Third, every successful wasm kernel compile rendered its whole module to text as an info diagnostic, accumulated throughyup::String +=on a String with no capacity reserve - quadratic in module size, so the largest example patch (TX81Z) effectively never finished compiling on wasm; the renderer now accumulates into a growth-friendlystd::string.
yup_dsp_jit: YDSP gains parameter smoothing. The newsmooth (x, tau)intrinsic is a one-pole ramp towardsxwith atau-second time constant, plus a snap on the first sample where the ramp cannot advance, so the target arrives exactly and leaves no denormal tail (the compare tests the step rather than the remaining distance: a float32 lerp stalls while still short of the target by roughlyulp / coeff, so any fixed epsilon would either be unreachable or truncate a fast ramp); it lives under the same restrictions as the delay primitives (per-sample body, outside loops, not in event handlers) and costs one hidden float slot plus one hidden int slot per call site (the coefficient'sexpis loop-invariant and hoisted out of the sample loop).[[ smoothing: <seconds> ]]on aninput value floatis sugar for it: one synthetic local is prepended to the per-sample body and that body's references to the endpoint are rewritten, so event handlers,funcbodies andgetParameter()keep seeing the raw target. Nothing changed in the IR ops, the codegen backends or the runtime - automation is still a step, the ramp happens in generated code. Every "YDSP Synths" example patch now smooths the parameters that actually step the signal -Delayfeedback/damping/mix,Chorusdepth/mix,Compressorthreshold/ratio/makeup,Distortiondrive/tone/mix,Reverbmix/damping/room size,Analog Sawresonance,FM Bellmod index,Wobble Leadboth cutoff extremes,Pulse Basswidth andElectric Piano's tremolo depth - so dragging a knob no longer zippers. Parameters that only set a rate (LFO rates, the compressor's attack/release) or that are read once in an event handler (envelope times, FM ratio, all the Electric Piano voice controls) are deliberately left stepped.Delay's time stays raw on a slew-rate argument: ramping an ~88000-sample read index over 20 ms scrubs the buffer at ~100 samples per sample, worse than the single jump it replaces - whereasChorus's depth spans only ~880 samples and so slews at roughly the rate its own LFO already does.Analog SawandPulse Basssmooth their cutoff at the filter coefficient rather than at the parameter -float k = smooth (1 - exp (...), 0.02)- because a parameter is sampled once per kernel invocation, so an expression built only from parameters is already hoisted out of the sample loop; smoothing the parameter itself would pull thatexpback in, while smoothing the coefficient keeps it hoisted and costs onelerp. The four voices' hand-rolledenvSmooth/smoothCoeffanti-click one-pole collapses to a singlesmooth (env, 0.003); not quite a pure refactor, as a voice's very first note now attacks slightly faster (the primed flag snaps the gain on the voice's first sample instead of ramping it up from zero). That is inaudible - the oscillator and filter state are zero at that point too, so the voice still starts from silence - and a stolen voice is unaffected, since state is not zeroed on steal and the smoother ramps exactly as before.yup_dsp_jit: YDSP gains four small language additions, all const-folded away before name resolution so nothing downstream of the semantic analyzer changes. Program-scopelet name = <constant expression>;declares a compile-time constant usable as astatearray size, aforbound or in any expression (imported constants are namespaced like imported processors, and redeclaring a constant's name is an error);statedeclarations take initialisers (state float feedback = 0.5;,state float table[32] = { ... };with trailing elements left zero) which are lowered into the processor'sinitkernel, synthesising one if the processor has no explicitinit { }block; the newsamplePeriodbuiltin is1 / sampleRate(loop-invariant, so the division is hoisted out of the sample loop); and[[ init: <value> ]]is accepted as an alias for a parameter's default value, so annotation blocks paste in unchanged (an explicit= <value>still wins). Constant folding now also handles binary arithmetic, so an endpoint default like= 1.0 / 3.0evaluates instead of silently becoming zero.yup_dsp_jit: YDSPforloop variables are now scoped to the loop body as documented - sibling loops may reuse the conventionalfor iwithout colliding, and the variable is no longer visible after the loop (previously the secondfor iin one body was rejected as a duplicate symbol).examples/graphics: new "Electric Piano" YDSP synth patch (data/synths/ElectricPiano.ydsp), a 16-voice additive electric piano - a 32-partial complex-rotation oscillator bank per voice with velocity-blended spectra, per-partial decay interpolated in 64-sample chunks, and a stereo triangle tremolo. The bank is stored as parallelstate float[32]arrays rather than a struct of complex values so the harmonic loop walks unit-stride. The two partial-weight spectra are hand-authored, not measured from a real instrument.yup_dsp_jit: YDSP is now a full MIDI/MPE instrument host. Processor-scope events grew fromnoteOn/noteOffto seven shapes namednoteOn(.pitch,.velocity,.isLegato),noteOff,pitchBend(.bendSemitones),pressure,slide,control(.control,.value) andprogramChange(.program) - driven by a single shape/field table that the analyzer, the IR builder and the runtime all consult, and lowered through two offset-carrying IR ops (loadEventFieldF/loadEventFieldI) that replace the previous per-field opcodes. Ingestion now goes throughyup::MPEInstrument, so plain MIDI and MPE share one path and note expression is keyed by note identity rather than by pitch: per-note bend/pressure/slide reach only the owning voice, sustain, sostenuto, reset-all-controllers and all-notes-off are honoured as ordinary note lifecycle, and all-sound-off (CC120) silences and re-runsinitat its exact sample offset. Voice behaviour is declared on the node -node v = Voice[8] [[ mode: poly, stealing: oldest ]],node b = Bass [[ mode: mono, priority: last ]]- with mono mode keeping an allocation-free held-note stack and flagging legato continuations through.isLegato. NewYdspAudioGraph::setMpeZoneLayout()/setLegacyMidiMode()(legacy is the default, so existing hosts are unaffected), andprepare()'s per-voice event capacity default rose from 32 to 64 for MPE traffic. Two behaviour changes: a second note-on at the same pitch and channel now retriggers (releases) the first instead of stacking a second voice, and expression for a note with no allocated voice is discarded and counted ingetDroppedEventCount().yup_dsp_jit: the native asmjit backend was split per architecture -YdspCodegenis nowYdspAsmJitCodegen(architecture-independent facade) with the shared lowering inYdspAsmJitCodegenImpland the per-target encodings inYdspAsmJitCodegenX64(x86-64 SSE) andYdspAsmJitCodegenARM64(AArch64 ASIMD) (backend/yup_YdspAsmJitCodegen{,.X64,.ARM64}.*); purely organizational, no behavior change beyond the class rename.yup_dsp_jit: new WebAssembly backend for emscripten/browser targets (previously the module hard-errored onYUP_WASM). Each kernel is lowered to a self-contained wasm module (backend/yup_YdspWasmEmitter.*binary writer +yup_YdspWasmCodegen), run by the browser's nativeWebAssemblyengine: the module imports the host's shared linear memory asenv.memory(host buffers addressed directly as i32 offsets, no marshalling), libm intrinsics come fromenvbacked byMath.*with C-exactround/copysign/fmod, and structured control flow lowers to wasm block/loop/if (zero-division-guardeddivI/modI,trunc-basedmodF). Kernels are instantiated per JS realm (main thread and audio-worklet thread each get their own copy, lazily on first use) and are keyed by unique per-kernel ids, so a worklet realm can never invoke an older graph's kernel with a newer graph's context when patches are swapped;YdspAudioGraph::prewarmKernels()pre-registers them in the calling realm, and the graphics example prewarms its graph from the audio callback. The asmjit dependency is now desktop-only, the wasm tests run under the emscripten node target, and theexamples/graphicsYDSP Synths demo works in the browser - its "Dump Asm" action prints the generated wasm as WebAssembly text (YdspWasmCodegen::toText, recorded as an info diagnostic at compile time, mirroring the asmjit assembly log on desktop). Defined-function indices are assigned in the wasm function index space (imported functions only; the memory import no longer shifts them).yup_dsp_jit: the whole module now usesyup::String/yup::StringRefinstead ofstd::string/std::string_view- AST fields, tokens, parser/analyzer/optimizer/codegen signatures and diagnostics messages - and reusesyup::Stringfacilities (getLargeIntValue(),getDoubleValue(),substring(),lastIndexOfChar(), concatenation) in place of the stdlib string helpers.yup_dsp_jit: node-level oversampling (node = P * N) now resamples throughyup::Oversampler(windowed-sinc, 2x/4x/8x) instead of the hand-rolled linear-interpolation upsample and boxcar-average downsample; the resampler introduces2 * SincRadius(16) samples of latency per oversampled node, oversampling requires equal input/output stream counts (or no inputs), and unsupported factors run the node at 1x.yup_dsp_jit:YdspAudioGraphnow exposes host-UI parameter metadata -getParameterCount()/getParameterInfo(slot)(qualified name,[[ name ]]display name, type, declared default,[[ min ]]/[[ max ]]bounds) plusgetInputStreamCount()/getOutputStreamCount()- so hosts can build sliders directly from a patch. Newexamples/graphics"YDSP Synths" demo (desktop-only, sinceyup_dsp_jitrequires asmjit) loads.ydspsynth patches fromdata/synths/, compiles them lazily withYdspCompiler, builds parameter sliders from the annotations, and drives them from a MIDI keyboard throughYdspAudioGraph::process(..., MidiBuffer*, ...).yup_dsp_jit: YDSP gains MIDI-driven events -input event noteOn/noteOffendpoints,event <endpoint> (<param>) { ... }handler blocks, andnode = Processor[N]voice banks with fixed-size, allocation-free voice stealing. MIDI byte-decoding and voice allocation live in the runtime (YdspAudioGraph); events and parameter automation are dispatched sample-accurately via runtime sub-block splitting (the same compiled kernel is re-invoked per sub-block - no codegen changes), which also powers a new stepped parameter-automation API (YdspAutomationEvent,YdspAudioGraph::getParameterSlot());YdspAudioGraph::processgains aMidiBuffer/automation overload andgetDroppedEventCount(). A graph-levelinput event midi;documents MIDI consumption; routing is structural (broadcast to every event-driven node). Fixed a pre-existing out-of-bounds param/meter indexing bug invalidateConnectivityfor stream-free, parameter-carrying processors.yup_dsp_jit: YDSP processors can now declarestructtypes with primitive and fixed-array fields, andstatevariables can be struct instances (Voice state;) or arrays of structs (Voice voices[32];). Fields are accessed with.(state.phase,voices[v].gate,state.buf[i],voices[v].buf[j]) and flatten into the existing scalar/array state slots - no IR or codegen changes required. Struct states are state-layout only (no struct values/locals/params yet).yup_dsp_jit: YDSP processors can now declare aninit { ... }block that runs once before audio starts (state writes, param reads, bounded loops; streams and delay primitives are rejected). Each init block compiles to a one-shot kernel sharing the processor's state layout;YdspAudioGraph::prepare()runs them in topological order, and the newYdspAudioGraph::reset()re-zeroes state and re-runs them.yup_dsp_jit: YDSP now supports the C-style bitwise operators&,|,^,<<,>>,~(and their compound assignment forms&=,|=,^=,<<=,>>=) onint32/int64operands, with C-compatible precedence (|<^<&< equality < relational < shift < additive), arithmetic (sign-propagating) right shift, and full support through the optimizer (constant folding + identity peepholes) and the asmjit backend on both x86-64 and AArch64.yup_dsp_jit: YDSP now supportsint32/int64andfloat32/float64primitive types end-to-end (float/intare aliases for the 32-bit types). Typing is strict: no implicit conversions, contextual literal adaptation only, with explicitint32(x)/int64(x)/float32(x)/float64(x)casts;if/&&/||/!requirebool, indices and loop bounds requireint32, and the delay primitives stayfloat32.yup_dsp_jit: float64/int64 streams, parameters and meters are supported through the whole pipeline: the asmjit backend emits SSE2 double instructions (movsd/addsd/comisd/cvtsd2ss/…) and 64-bit GPR arithmetic on x86-64, register-sizedfadd d/fcvt/sxtw/scvtfon AArch64, and width-matched libm variants (sinvssinf).YdspAudioGraph::processgained per-stream/param element-type introspection and typed parameter accessors (getDoubleParameter/setIntParameter/…).yup_dsp_jit: internal module sources were reorganized for maintainability - the semantic analyzer, the IR optimizer and the graph runtime monoliths were split into single-responsibility files (analysis/semantic/,optimiser/builder/+optimiser/passes/,runtime/graph/); purely organizational, no API or behavior change.examples/graphics: the "Analog Saw" YDSP synth patch now uses a polyBLEP band-limited sawtooth (naive ramp minus the polyBLEP residual at the phase wrap) instead of the naive aliasing ramp, demonstrated through thepolyBlepfunction.examples/graphics: every oscillator left in the YDSP synth patches underdata/synths/is now band-limited too - the remaining naive saws (Haas Widener, Parallel Drive, Parallel Rack, Wobble Lead) and the Pulse Bass pulse use the same polyBLEP treatment (two edges for the pulse), and the FM Bell is rebuilt as an additive (Le Brun) expansion whose Bessel-weighted sidebands are silenced above Nyquist, so none of the demo voices alias.examples/graphics: new "Wave Lab" YDSP synth patch (data/synths/WaveLab.ydsp) demonstrates four polyBLEP band-limited oscillators behind onewaveselector knob - saw, square, triangle (whose corners are rounded by the exact antiderivative of the BLEP residual) and a width-morphable pulse.examples/graphics: the "YDSP Synths" demo gains a "Dump Asm" button that prints the asmjit assembly listing of the current patch's compiled kernels to the console (replayed from the graph's compile-time diagnostics).examples/graphics: the "YDSP Synths" demo gains Performance/Editor tabs - Performance keeps the existing knob grid, keyboard and oscilloscope; Editor is a liveCodeEditor(YDSP syntax highlighting) bound to the selected patch's source with a Compile button that recompiles it throughYdspCompilerand swaps in the resultingYdspAudioGraphon success (refreshing the Performance tab's knobs), or shows the compiler's diagnostics (source line + caret) in a red panel on failure.
yup_audio_basics: newAudioLockTypealias -yup::RecursiveSpinLockon wasm andyup::CriticalSectionelsewhere - used by the module's audio types (MidiKeyboardState,UMPKeyboardState,Synthesiser,MPESynthesiser*,MPEInstrument, and the audio sources). On wasm aCriticalSectioncan block in afutex_wait, which is fatal on the browser audio-worklet thread;RecursiveSpinLockbusy-waits on an atomic and never blocks, while keeping the re-entrancy ofCriticalSection.yup_audio_basics:MPEInstrumentcan now be driven from an audio callback without allocating - newreserveNotes (int)preallocates the note-tracking storage, and note removal uses theArray::remove()(inyup_core) which is now not allocating on removal.
- New
asmjit_librarymodule (thirdparty/asmjit_library): AsmJit machine-code generation library (core, x86, AArch64 and ujit backends), statically linked via the YUP module system.
Componenttransforms now apply hierarchically: a parent's transform carries its children when painting, hit-testing, delivering mouse and drag-and-drop events, and inlocalToScreen()/screenToLocal()and the helpers built on them. Insidepaint(),Graphics::getTransform()holds the linear part of the composed transform and the drawing area position the translation, so untransformed hierarchies paint exactly as before. The component transform is no longer applied twice to effect and cached composites.yup_dsp_jit: the public API now usesyup::String/yup::StringRefinstead ofstd::string/std::string_view-YdspCompiler::compile, everyYdspAudioGraphparam/meter accessor,YdspDiagnostics::addError/addWarning/addInfo/setSource/toString,YdspDiagnostic::message,YdspKernelReport::name/loopBounds(nowyup::StringArray) andYdspParameterInfo::name/displayName. Callers passingconst char*orstd::stringkeep working unchanged (implicitStringRefconversion); explicitstd::string_viewarguments must switch toyup::StringRef.yup_dsp_jit:YdspAudioGraph::process()no longer takes rawvoid*stream pointer tables. Stream buffers are now passed as typed spans -process (yup::Span<const YdspInputBuffer>, yup::Span<YdspOutputBuffer>, int)- whereYdspInputBuffer/YdspOutputBufferarestd::variants ofyup::Span<float|double|int32_t|int64_t>. The active variant alternative carries the buffer's element type, so a mismatched buffer is ignored and reported through the returnedYdspProcessResultvalue instead of being reinterpreted; the generated-kernel ABI (YdspKernelContext) is unchanged. Theprocess32()convenience overload was removed - callers pass typed spans toprocess()directly.- macOS: OpenGL rendering backend removed in favor of Metal only
LottieReader::parseFile(),parseData(),parseStream(), andparseFromZip()now returnResultValue<AnimationComposition::Ptr>and no longer take a trailingString* outErrorout-parameter; checkwasOk()/failed()and read the message viagetErrorMessage().AnimationFrameExporteris now an instance-based class bound to aGraphicsContext(constructAnimationFrameExporter exporter (ctx);then callexporter.renderFrame(anim, …)/exporter.renderAllFrames(…)/exporter.exportToGif(anim, …)), so it can own and reuse the GPU matte-composite pipeline across frames instead of recompiling it per frame. TheexportToGif(frames, frameRate, …)frame-sequence encoder remains a static helper.- Config macro
YUP_EMBED_DEFAULT_THEME_TEXT_FONTrenamed toYUP_EMBED_DEFAULT_THEME_TEXT_SERIF_FONT; the embedded default text font now only covers the serif font. A newYUP_EMBED_DEFAULT_THEME_TEXT_MONOSPACE_FONTconfig selects whether the monospace theme font is embedded. SyntaxDefinitionno longer carries colors:getColor(),getSelectionColor()and the JSON"colors"section were removed. Token and editor colors now live in the newCodeEditorScheme(seeCodeEditor::setScheme).Artboardno longer defines its ownLayoutandAlignmentenums; it uses YUP's sharedFittingandJustificationtypes.setLayout/getLayoutbecamesetFitting/getFittingtaking astd::optional<Fitting>, andsetAlignment/getAlignmentbecamesetJustification/getJustificationtaking aJustification. The mapping isfill→Fitting::fill,contain→Fitting::scaleToFit,cover→Fitting::scaleToFill,fitWidth/fitHeightunchanged,none→Fitting::none,scaleDown→Fitting::centerInside, andlayout→std::nullopt(no fitting of our own, so the artboard resolves its authored Rive layout constraints).Alignmentmaps 1:1 onto theJustificationcorner/edge names, withtopCenter,bottomCenterspelledcenterTop,centerBottom.Fittingvalues Rive has no equivalent for (tile,centerCrop,stretchWidth,stretchHeight) are accepted and fall back toscaleToFit; becauseJustificationis a bitfield it also admits partial states the old enum could not express, and an axis with no flag set is centered on that axis.NodeAttachmentOptions::pivot/anchoralready took aJustification, so the class now speaks one alignment vocabulary throughout.Fontloading is now static-only: the instanceloadFromData()/loadFromFile()methods were removed in favor ofFont::loadFontFromData(),Font::loadFontFromFile(),Font::loadFontFromFirstAvailableFile(),Font::loadSerifSystemTextFont()andFont::loadMonospaceSystemTextFont(), all returningResultValue<Font>.- The
yup_rhidescriptors now own their data instead of pointing at caller-managed storage.GpuVertexBufferLayout::attributesandGpuPipelineOptions::vertexBuffersarestd::vector(droppingattributeCountandvertexBufferCount),GpuPipelineOptions::colorTargetsis astd::vectorcapped at four (droppingcolorTargetCount),GpuShaderSource'scode/bindingMap/glFixupare nowstd::vector<uint8>(owning) instead ofSpan<const uint8>withentryPointaString, andGpuTextureDesc::label/GpuSamplerDesc::labelbecameString. Call sites that built astatic constexprattribute table to keep it alive can now build the layout inline; aoptions.colorTargets[0].format = …becomesoptions.colorTargets.emplace_back().format = …. UsegpuShaderSourceBytes()to build aGpuShaderSource::codeblob from source text.GpuDevice::beginOffscreennow takes the newGpuFrameDescriptorinstead ofrive::gpu::RenderContext::FrameDescriptor. GpuDevice's move constructor and move assignment are now= delete. They were publicly defaulted on aReferenceCountedObject, so moving a device out from under liveGpuDevice::Ptrholders corrupted the refcount.GpuBuffer::Impl,GpuBuffer::getImpl()andGpuBuffer::createWithImpl()moved frompublic:toprivate:(they were marked@internalby comment only); the backend factories that use them are friends.GpuTexture::getOreTexture()andGpuTexture::getRenderImage()were removed - neither had any caller.DragAndDropDatais now a MIME store and moved fromcomponent/to the newdragdrop/folder. Payloads are held as anArray<ClipboardData>under well-known MIME types (DragAndDropData::mimeTypeText/mimeTypeUriList/mimeTypePng), with text, files, URIs and images exposed as convenience accessors over that single store plus an optional same-processvarnative object. ConsequentlygetFiles(),getText()andgetUris()now return by value (Array<File>/String/StringArray) rather thanconst&(they decode from the MIME store), and the class gainedwithImage/withMimeData/withNativeObjectwith the matchingget*/has*/getMimeTypes/getAllMimeDataaccessors. There are no lazy or promised data providers: every MIME blob is an eagerly-ownedMemoryBlock.- The five drag-and-drop virtuals on
Component(isInterestedInDrag,itemsDropped,itemDragEnter,itemDragMove,itemDragExit) and theComponent::internalItemDrag*dispatch they fed have been removed, soComponentno longer carries any drag-and-drop surface. Drop targets are now an opt-in mixin: derive fromDragAndDropTarget(indragdrop/) alongsideComponentand overrideisInterestedInDragSource/itemDropped/itemDragEnter/itemDragMove/itemDragExit— or assign the matchingstd::functionmembers — each receiving a singleDragAndDropSourceDetailsthat carries the payload, the source component, the target-local position, the allowed actions and the suggested action. The library finds targets with adynamic_cast<DragAndDropTarget*>(seeDragAndDropTarget::dispatchItemDrop()and friends), preserving the previous enter/move/exit and child-to-parent drop-bubbling semantics. The Python bindings for the removedComponenthooks were dropped;DragAndDropTargetandDragAndDropTargetComponentare bound, as areDragAndDropSourceand itsDragOptions. Oversamplerwas renamedSincOversampler(resampling/yup_SincOversampler.h) now that it is one of two oversampler designs, and theOversampler2xFloat…Oversampler8xDoublealiases were removed:HalfbandOversampler2xFloat…HalfbandOversampler32xDoublenameHalfbandOversamplerinstantiations with the default design (100 dB, passband to 0.45 of the input rate, stopband from Nyquist, linear-phase FIR). The call surface is identical, butgetLatencyInSamples()reports the cascade's own delay (about 140 input samples for the 4× round trip with the defaults, against 32 for the radius-16 sinc) andgetGenerationLatencyInSamples()is no longerstatic constexpr. Code that needs the sinc design or a non power-of-two factor should spell outSincOversampler<SampleType, Factor, Radius>ComponentNative::setFocusedComponent()takes a secondFocusChangeTypeargument saying what moved the focus. It defaults toFocusChangeType::focusChangedDirectly, so callers are unaffected, but any class implementing the pure virtual has to match the new signature.
- Added
File::bundleDirectoryandFile::hostBundleDirectoryspecial locations: the resource root bundled with the current YUP binary itself (a plugin's own.vst3/.component/.clap, an app's own bundle, or an Android APK's assets) versus the host application's bundle when running as a plugin. On Android, files underbundleDirectoryare read directly out of the APK viaAAssetManagerthrough the normalFile/FileInputStreamAPI - no first-run copy to disk - and are read-only (createOutputStream(),deleteFile(),createDirectory()fail cleanly). Files can be placed there with the newyup_add_bundled_resources()CMake function - Added a
YAMLclass: a self-contained YAML parser and writer converting between YAML text andvar(parse,fromString,toString,writeToStream,FormatOptions), with core-schema type resolution, block/flow collections, block scalars, and anchors/aliases/merge keys - Added a
CancelTokenclass (threads/yup_CancelToken.h): a thread-safe, copyable observer token withwasCancelled(), blocking observation viawaitForCancellation(), and callback observation viaregisterCallback()/Registration - Added a
CancelTokenSourceclass (threads/yup_CancelTokenSource.h): a move-only RAII owner of aCancelTokenthat is the sole canceller, requesting cancellation automatically when destroyed (unless moved-from), with observer copies obtained viagetToken() - Fixed
IPAddressdisagreeing with itself about byte order inside each 16-bit group. The class reinterprets its byte array asuint16groups on little-endian hosts, but the string parser packed a parsed group the opposite way round, so an address parsed from"fe80::1"did not equal the same address built from its integer groups, and an IPv4-mapped address did not round trip throughtoString()and back or compare equal to its IPv4 counterpart.toString(), bothuint16constructors, the parser andconvertIPv4MappedAddressToIPv4()now pack and unpack a group through the same low-byte-first layout WaitableTimernow meets its deadline on Apple platforms with a short series of absolutemach_wait_until()sleeps, each covering 75% of the time still left, followed by a 50us busy-wait. Darwin coalesces timer expiries into a window proportional to the requested sleep duration, so the previous condition-variable wait overshot in proportion to how long it was asked to wait, whereas shrinking sleeps make that error decay geometrically over a handful of syscalls. The condition-variable fallback is no longer compiled on AppleWaitableTimerruns the same cascade on Linux and Android using absoluteclock_nanosleep()sleeps, and asks for the smallest possible timer slack on the calling thread the first time it waits (Linux applies 50us of slack to every timer expiry outside the realtime scheduling classes, which an application is not normally allowed to join). The busy-wait budget is 150us there rather than Apple's 50us, measured on a Galaxy S25 running Android 16: 50us leaves a p99 wake error of 180us, 150us brings it to 86us, and going higher only costs more CPU. The condition-variable fallback now only remains for Windows (when the waitable timer cannot be armed) and WebAssembly- Fixed
WaitableTimer::waitUntil()returning immediately for every deadline within one millisecond. The platform-specific final wait now runs for short deadlines instead of starting time-sensitive work early - Fixed
Thread::startRealtimeThread()leaving the thread unstartable on macOS/iOS when the realtime upgrade failed:createNativeThread()reported the failure but kept the handle of the thread that had already bailed out, soisThreadRunning()stayed true and any laterstartThread()was refused
-
Added
HalfbandOversampler(yup_dsp/resampling/), a power-of-two oversampler built from a cascade of 2× halfband stages so that only the stage next to the base rate has to be steep.Designselects the family and targets:linearPhaseFIRdesigns Kaiser-windowed halfbands, verifies every stage's stopband numerically atprepare()and keeps both latencies whole input samples by rounding the cascade's delay at the top rate;polyphaseIIRdesigns elliptic halfbands as two allpass branches (Valenzuela & Constantinides) for a few multiplies per sample and minimal, frequency-dependent latency. Defaults are 100 dB of rejection, a passband to 0.45 of the input rate and, for the FIR, a stopband starting at the input Nyquist, so nothing folds back into the band (a pure halfband first stage, selected withstopbandEdge = 1 - passbandEdge, costs a quarter as much but folds 0.5 to 0.55 of the input rate into the top of the band, which is what the IIR family always does). With the defaults 4× decimates in about 300 MACs per input sample against 128 for the radius-16SincOversamplerat 90 dB, a 0.36 passband and a 40 dB leak at Nyquist; at 32× it is about 600 against 1024, and the IIR cascade needs about 76 multiplies. The public methods mirrorSincOversamplerso the two are drop-in replacements.ModulatedOscillatornow decimates through it: itsSincRadiustemplate parameter is gone (ModulatedOscillator<SampleType, OversampleFactor, CoeffType>),prepare()takes an optionalHalfbandOversamplerDesign, andgetLatencyInSamples()became an instance method that reports the design's latency afterprepare() -
SincOversampler(the single-stage polyphase sinc formerly namedOversampler) is several times faster and rejects images and aliases far better. Each channel now keeps its history contiguously in front of the staging buffer so every output sample is onedotProduct(SIMD forfloat/floatand, newly,double/doubleviaFloatVectorOperations), replacing the per-tap circular-buffer modulo and strided table reads; the kernels are prebuilt phase-major with the gain baked in. Both kernels use Kaiser β = 9 (was 5), andSincTable::applyKaiserWindownow spans exactly the kernel radius (it previously windowed over(SincRadius + 1) · OversampleFactorentries while the kernel was used out toSincRadius · OversampleFactor, cutting it off at about 8 % of the window peak and capping the rejection ofSincOversamplerandResampleralike). TheOversampler2xFloat…Oversampler8xDoublealiases moved from radius 8 to radius 16, so their reported latency grows from 16 to 32 input samples, andOversampler16xFloat/Oversampler32xFloatplus theirDoublevariants were added -
Added
PrismSpectrum(yup_dsp/oscillators/), promoting the graphics example's Prism recipe into the library. It shapes aFourierSeriesthrough a raised-cosine ridge comb in log2-harmonic space, magnitude companding, exact spectral pulse-width modulation and a quadratic phase dispersion, renormalizing to the source'ssum |c|. A spectral tilt, an odd/even balance, a movable formant resonance and a pseudo-random phase scatter sit alongside those, each continuous through its own neutral value so it can be modulated without a step; scatter shares dispersion's rotation and so costs nothing extra. Every stage is a per-harmonic scale or rotation, so none can add a frequency and the backend's own bandlimiting is preserved. There is no transform behind it, so a shape can be re-derived once per audio block and the controls modulated;prepare()precomputes the per-harmoniclog2tables. The pulse-width depth fades in over the firstsqueezeFadeWidth, because the raw factor tends to a differentiator rather than to unity as the width falls, which would otherwise make zero a discontinuity -
Added
VAStateVariableFilter(yup_dsp/filters/), a topology preserving transform state variable filter modeled on Zavalishin's papers. One pass yields lowpass, highpass, constant-skirt and unity-gain bandpass, band shelf, notch, allpass and lowpass-minus-highpass outputs; resonance can be set in 0..1 or as Q, the shelf gain in dB, and the complex response is analytic -
Added
LFO(yup_dsp/oscillators/), an allocation-free control-rate oscillator with sine, triangle, sawtooth, square and seeded sample-and-hold shapes, a phase offset, andskip()for block-rate consumers -
Added
WaveformBank::refreshFrames(), replacing a prepared bank's coefficients without allocating and keeping the per-level harmonic limitsprepare()chose, plusModulatedOscillator::setBank()for adopting a bank rebuilt off-thread -
ModulatedOscillator::Parametersmoved to namespace scope asModulatedOscillatorParameters(still aliased as the nestedParameters), and the modulation maths moved intodetail::ModulatedOscillatorVoiceso oscillators can compose over it without duplicating the phase map, sync and BLEP/BLAMP corrections -
Added a waveform selector (sine, triangle, saw, square) and 16× / 32× sweep oversampling modes to the spectrum analyzer example; the non-sine shapes are naive, so the sweep oversampling modes show their aliasing suppression. The oversampled sweeps now generate at a multiple of the device rate and decimate straight to it through
SincOversampler::beginGeneration()instead of passing through the radius-8 resampler, whose 54 dB stopband was setting the alias floor regardless of the oversampling factor -
Reduced the graphics synthesizer example to a single Prism oscillator per slot, deriving each slot's spectrum once per block rather than once per voice so the shape controls can be modulated, and exposing squeeze, squash, tilt, odd/even, formant and scatter alongside ridges, color and dispersion. The sync modes moved onto Prism: the shaped series goes through
SyncSpectralResamplerat the slot, driven by the sync mode chooser and a new sync ratio knob, so unison satellites and the waveform preview follow it for free; the Modulated algorithm, its FM knobs and the algorithm chooser are gone from the example. The example then grew a per-voiceVAStateVariableFilterstage with drive and keytracking, a second envelope, two per-voiceLFOs (retriggered per note or free running) with displays following the newest voice, and an eight-slot modulation matrix on a second page: every source is per voice, and a voice derives a private spectrum only while a route reaches one of its spectrum parameters. Knob captions show the value while dragging, and Randomize now covers every voice property (oscillators, filter, envelopes, LFOs and a few live routes) while leaving volume, voice mode, glide and MIDI input alone. The example is split intoexamples/graphics/source/examples/audio/(settings, engine, panels, modulation page) -
Improved the graphics synthesizer example with a spectral Prism oscillator, octave and cents controls, poly/mono/legato modes, portamento, rendered voice-stealing tails, a revised instrument layout, and audio-load/overrun metering. Reduced callback and scope overhead, removed output waveshaping, and added regression coverage.
-
Added shared
WaveformBankand oversampledModulatedOscillatorwith through-zero FM, PM, phase distortion and fractional hard sync. Added direct generation toSincOversampler; fused spectral SIMD accumulation, amortized additive phasor trigonometry, and corrected sync bandwidth refresh and Nyquist boundaries. -
Fixed the pulsar spectral resampler's alternating coefficient sign and corrected oscillator regression tests for fixed-size copies, startup crossfades, and spectral window leakage.
-
Added a
TuningMapclass (midi/yup_TuningMap.h): maps MIDI note numbers to frequencies under an arbitrary scale and key map, loading Scala.sclscale files and.kbmkey map files vialoadScale()/loadKeyMap()(which return ayup::Resultand keep the previous tuning when a file fails to parse).isNoteActive()reports the notes a key map asks to retune, taken from the range in its header unless the file carries< first lastlines, which declare it instead -
FFTProcessoris now templated on the sample type -FFTProcessor<float>(the default) orFFTProcessor<double>- and every backend (PFFFT, Apple vDSP, Intel IPP, FFTW3 and the Ooura fallback, which now ships both afloatand adoubleimplementation) gained a native double-precision path. References to the nested scaling enum need qualifying, e.g.FFTProcessor<float>::FFTScaling::asymmetric -
Fixed the PFFFT-backed
FFTProcessorrequiring its input and output buffers to be SIMD aligned: the transforms are now staged through buffers owned by the PFFFT backend and allocated with PFFFT's own aligned allocator, so the public API accepts buffers with any alignment (the double-precision real transform previously hit PFFFT'sVALIGNEDassertion when handed a plainstd::vector<double>) -
Fixed the double-precision Ooura FFT translation unit only building on GCC/Clang: it declared every internal helper (
makewt,cftfsub,bitrv2, ...) inside the body of the functions that call them, and a block-scope declaration insidenamespace yupdeclares a global function, soyup::cdftreferenced a::makewtthat no one defined and the Windows link failed with 30 unresolved externals. The declarations now sit at namespace scope, matchingyup_OouraFFT8g_float.cpp -
Added the
yup_dsposcillator classes (oscillators/yup_FourierSeries.h,oscillators/yup_SyncSpectralResampler.h,oscillators/yup_AdditiveOscillator.h,oscillators/yup_WavetableOscillator.h,oscillators/yup_SyncOscillator.h), an alias-free synchronizing oscillator following Roth, Keller, Castaneda and Studer, "Alias-Free Oscillator Synchronization via Additive Synthesis" (DAFx26, paper 49):FourierSeriesholds the coefficients and the waveform presets,SyncSpectralResamplerrewrites them from a follower period ratio into hard, mirrored or pulsar sync,AdditiveOscillatorandWavetableOscillatorsynthesize the result using only the harmonics below Nyquist, and theSyncOscillatorfacade offers both backends and refreshes them from a once-per-blockupdate().utilities/yup_DspMath.halso gainedfillHarmonicPhasors
-
Added
Vector3,Matrix4(column-major, with perspective / orthographic / look-at factories andrive::Mat4converters) andRay(viewport unprojection, plane and triangle intersection), plusMeshSurfaceMapperin the newmeshes/folder, mapping between viewport points and texture coordinates of any triangle mesh. -
SVG drawables now prepare text layout, fitted image bounds, gradients, local dashed paths, marker placements and path bounds while parsing, avoiding repeated CPU work and failed image-resolution retries during rendering.
-
Lottie loading now prewarms static shape, mask, gradient and stroke-dash caches plus spatial motion-path lengths, so their first rendered frame does not perform one-time geometry or paint preparation; static gradients and masks are reused on subsequent frames.
-
Image::getWidth()andImage::getHeight()now return 0 on an invalid image instead of asserting and dereferencing null. Other accessors and pixel access still assert, as documented -
GpuTexturenow caches the backend texture views it hands out, keyed by view descriptor. A render pass previously allocated a freshore::TextureViewfor every attachment on every pass and for every sampled texture on every draw - all identical frame after frame - which on Metal made building the attachment descriptors cost more than creating the command encoder they were for. The cache needs no invalidation because aGpuTexturewraps one underlying texture for its whole lifetime:GpuCanvasandGpuTargetbuild a newGpuTexturewhenever their backing changes -
GpuRenderPassnow encodes every draw of a pass into a single backend render pass instead of opening and closing a fresh one per draw. Eachdraw()/drawIndexed()previously created its own command encoder - two per frame in a simple scene, one per pipeline in a multi-pass effect - which on Metal maderenderCommandEncoderWithDescriptor:one of the most expensive calls in a frame. The pass is opened lazily by the first draw and registered with the context, so beginning another one still auto-closes it, andGpuFrame::submit()closes a pass the caller left open before committing. Attachments must now be bound before the first draw, which is asserted -
GpuFrameno longer blocks on the GPU when it goes out of scope. Ending a frame cost a full pipeline stall on every rendered frame, justified by the encoded passes holding raw pointers to the frame's transient resources - but the backends YUP ships construct their ore context with noGPUResourceManager, so post-submit lifetime is already handled by Metal's command-buffer retention, OpenGL's deferred name deletion, and D3D11/WebGPU reference counting. A frame now hands its texture views, samplers and pooled uniform buffers to the device tagged with a generation, and a generation is released once enough later frames have begun.GpuFrame::waitForGPU()is unchanged and remains the explicit opt-in before a CPU readback -
GpuFramenow reports realcurrentFrameNumber/safeFrameNumbervalues to the ore context, which previously received a value-initialised descriptor leaving both at zero on every frame - so a manager-backed backend (ore's Vulkan and D3D12 paths) could never reclaim anything -
Fixed the Metal
GraphicsContextcreating its ownMTLCommandQueueinstead of sharing theGpuDevice's, even though the device already exposedgetDevice()/getCommandQueue()for that purpose. Metal orders work within a queue but never between two of them, so present work encoded on the private queue could run before the offscreen RHI work it samples had finished. Every other backend already shares a single command stream (D3D11 shares the immediate context, WebGPU sharesdevice.GetQueue(), OpenGL has one current context) -
Added
GpuTexture::create()andGpuTexture::upload(), so textures can finally be allocated and filled directly instead of only being obtained from aGpuTarget/GpuCanvas. Covers every shape the backend layer supports - 2D, cube, 3D and 2D-array - with mip levels, MSAA sample counts and the full 24-format list (8-bitr/rg/rgba/bgra, 16- and 32-bit float,rgb10a2unorm,r11g11b10float, the depth/stencil formats and the BC/ETC2/ASTC block formats) -
Added
GpuSamplerandGpuRenderPass::setSampler(): filtering, wrap mode per axis, mip filter, LOD clamp, anisotropy and comparison samplers. Slots the caller does not set keep the previous hardcoded linear / clamp-to-edge sampler, so existing pipelines are unchanged. A comparison-sampler binding now also gets a comparison function, without which the binding was invalid -
Added
GpuTarget::create (device, GpuTextureDesc)andGpuTarget::createFromTexture (device, texture, view), which give render-to-mip and render-to-cube-face. Thecreate (device, width, height)overload still goes through the Rive 2D canvas allocator and is unchanged. Note that a mip- or layer-narrowed view is honoured for attachments on every backend but not for sampling on OpenGL / OpenGL ES, which falls back to the base texture - read a specific mip with an explicittextureLod()and a LOD-clampedGpuSamplerinstead -
Added
GpuDevice::isFormatSupported(),isFormatRenderable(),isAnisotropicFilteringAvailable()andgetMaximumSampleCount(). Float colour attachments are extension-gated on OpenGL ES / WebGL2, so a caller can now degrade instead of rendering black -
GpuRenderPassnow binds the attachmentsGpuPipelineOptionswas always able to describe: up to four colour attachments viasetColorAttachment()(colorTargets[4]previously documented MRT but the pass hardcoded a single attachment), a depth/stencil attachment viasetDepthStencilAttachment()(depthStencil.enabledpreviously compiled a pipeline that no pass could satisfy), and an MSAA resolve destination viasetResolveTarget() -
GpuRenderOptionsgained explicitGpuLoadOp/GpuStoreOpper attachment, plusGpuDepthStencilOptionsfor depth and stencil. The{ bool clear, GpuColor }form is kept as a shorthand. Attachments are now cleared before the first draw of a pass only, so several draws can accumulate into one surface and share one depth buffer instead of each draw wiping its predecessors -
GpuRenderPassbind groups now only carry the slots the bound pipeline's layout actually declares. Binding state outlives a single draw, so a slot bound for an earlier draw with a different pipeline used to be handed to the backend anyway, which rejects the whole group - silently dropping every binding in it. A bind group that still fails to build is now reported instead of leaving the draw reading whatever was bound before it -
GpuRenderPass::draw()anddrawIndexed()now forwardinstanceCount,firstVertex,firstInstance,firstIndexandbaseVertex, so instanced drawing works -GpuVertexStepMode::instancewas previously declarable but could never advance past instance 0. AddedsetViewport(),setScissorRect(),setStencilReference()andsetBlendColor(), all sticky across the draws of a pass -
GpuPipelineOptionsgainedGpuColorTarget::writeMask(GpuColorWriteMask), the remaining ore blend factors (srcAlphaSaturated,blendColor,oneMinusBlendColor) and the remaining vertex formats (sint8x4,uint16x2,sint16x2,unorm16x2,snorm16x2,uint16x4,sint16x4,float16x2,float16x4,uint32) -
Fixed shader binding maps hardcoding every texture binding as a non-multisampled 2D float texture: the reflected image dimension, arrayed flag, depth flag and multisampled flag are now carried through, so a
textureCube,texture3D,texture2DArrayor depth texture no longer fails bind-group-layout validation on WebGPU.backendSpaceis also now set to the binding's group, which was wrong for anyset != 0 -
Fixed
GpuSamplerDesc::labeldangling:GpuSampler::create()copied the rawconst char*into a retained member and handed it back through the publicgetDescription(), so any caller passing a temporary got a dangling pointer for the sampler's lifetime. Both it andGpuTextureDesc::labelare nowString -
Fixed
GpuPipelineCachenot hashingGpuColorTarget::writeMask, so two pipelines differing only in write mask collided on the same cache key even though the mask is plumbed into the pipeline -
Fixed the pipeline entry-point names being borrowed rather than owned:
ore::Pipelinekeeps a shallow copy of thePipelineDescand dereferences its entry-point strings at draw time, whilecompileFromBundle()pointed them at a localString. The compiledGpuPipelinenow owns them, alongside the vertex layouts it already copied -
Fixed the WebGPU compute path ignoring the shader source length and requiring NUL-terminated code, which is not what an RSTB blob is
-
A render pass that binds a pipeline incompatible with its attachments (colour format, sample count or depth presence) now asserts and logs the backend's own diagnostic, instead of silently rendering nothing
-
Added a
PBR IBLexample to the graphics demo: a procedural sky baked into a float cube map face by face, convolved into an irradiance cube, prefiltered into a roughness mip chain, plus a split-sum BRDF lookup table and a CPU-uploaded albedo / normal map, shading an instanced grid of spheres against a real depth buffer -
Fixed
GpuDevice::isComputeAvailable()reporting from a runtime GL version probe on WASM / WebGL, where no GL compute implementation is compiled in at all; its documentation also claimed compute was available on D3D12 and Vulkan (neither backend exists) and unavailable on OpenGL (where it is implemented) -
Fixed
native/yup_GpuDevice_dawn.cppguarding its whole body onRIVE_DAWNwhile the module includes it underYUP_RIVE_USE_DAWN, so the Dawn device could compile to nothing -
Added
GpuFrameDescriptor(rhi/yup_GpuTypes.h), a field-for-field mirror ofrive::gpu::RenderContext::FrameDescriptorthat keepsGpuDevice::beginOffscreen's public signature free of a Rive type.GpuCanvas::beginDraw()gained an optionalconst GpuFrameDescriptor¶meter (defaulting to today's behaviour: clear to transparent black, no msaa), giving callers control overmsaaSampleCount,ditherMode(newGpuDitherModeenum),loadOpandclearColorfor the offscreen 2D frame it opens -
GpuCanvas::beginDraw()and the offscreenGraphicsconstructor take ascale(canvas pixels per logical unit):getContextScale()reports it and the default drawing area is the canvas size divided by it. -
Fixed GPU compute silently stalling after a few frames on OpenGL with some drivers (AMD desktop GL): compute now runs on a dedicated, unshared GL context (
GpuDevice::Options::computeContextActivator, routed throughGpuDevice::runOnComputeContext()) which exclusively owns every compute resource — pipeline compilation, dispatches, storage buffer create/update/readback and deletion — falling back to the rendering context when unavailable. The GL compute pass also saves and restores the program andGL_UNIFORM_BUFFERbindings it touches, so it can no longer desync Rive's cached GL state -
Fixed GL storage buffers being deleted right after creation (moving a
GpuBuffer::Implcopied the plain GL buffer name, so the moved-from object's destructor freed the just-created buffer) and a crash when releasing GPU buffers or compute pipelines after their window closed (GL releases are routed through the owning device and skipped once the window's contexts are gone) -
Fixed Emscripten randomly rendering nothing or freezing the tab: the render-thread rework unbound the GL context between frames so the render thread could take it, but on Emscripten rendering is timer-driven on the single browser thread and message-thread work between frames (image decodes, font atlas uploads) still issues GL calls — with no context current those throw in JS and kill the main loop. With timer-driven rendering the window context now stays permanently current
-
Fixed Emscripten freezing the browser tab when a demo requested a headless GL compute device: with no current WebGL context every GL call throws in JS, killing the requestAnimationFrame main loop. The GL
GpuDevicenow fails construction gracefully when no WebGL context is current andisComputeAvailable()reports false on a device whose GL initialization failed.GpuAudioProcessingDemoalso sized its CPU ring buffers only when GPU compute was available, so the no-GPU audio path wrote and read out of bounds -
Fixed
SDLComponentNative::renderFrame()on Emscripten encoding a frame with a zero-sized render target, which trips Rive'sbeginFrame()assertion — fatal there, since an assertion abort throws inside therequestAnimationFramecallback and permanently kills the browser's main loop. It now skips rendering entirely until a non-zero content size has been observed -
Added a native WebGPU
GraphicsContextbackend for Emscripten via the Emdawnwebgpu port (RIVE_WEBGPU=2+--use-port=emdawnwebgpu, enabled with theENABLE_EMSCRIPTEN_WEBGPUparameter ofyup_standalone_app), rendering Rive content through the browser's WebGPU API without Dawn -
Fixed
GpuFrame::begin()aborting on the Emscripten WebGPU backend: the WGPU context now creates and submits its own command encoder when no external one is provided, matching the Metal/GL/D3D11 self-managed frame model -
Fixed a crash on Windows when creating any native window: the D3D11
GpuDevicewas built with an already moved-fromID3D11Device, and the Direct3DGraphicsContextcreated a second device whose swapchain textures could not be used by the render context. Both now share a singleID3D11Device -
Fixed the Emscripten WebGPU
GraphicsContextnever storing its surface size, leaving the offscreen copy at 0x0 -
Fixed
GpuDevice::updateBuffer()failing for every vertex, index and uniform buffer on the WebGPU, Dawn and D3D11 backends: those overrides handled native storage buffers only and returned false instead of delegating ore-backed buffers to the base class, the way the Metal and OpenGL overrides do -
Implemented
GpuDevice::readBuffer()for D3D11, which previously reportedisComputeAvailable()but had no override, so every storage buffer readback silently failed through the base class. It copies into a cachedD3D11_USAGE_STAGINGbuffer on the immediate context (ordered after the dispatch) and maps it for reading -
GpuComputePasson D3D11 now unbinds the UAV slots it bound when the pass finishes, so a storage buffer is no longer left bound for writing while a later readback or draw reads it -
Fixed
GpuDevice::readBuffer()never succeeding on the Emscripten WebGPU backend: it mapped its staging buffer withWGPUCallbackMode_AllowProcessEventsand then tested the result in the same call, but WebGPU buffer mapping only resolves through the JavaScript event loop, so the callback could not have run. The WGPU backend now pipelines the readback over a ring of three staging buffers usingWGPUCallbackMode_AllowSpontaneous, which completes on its own between main-loop ticks - no ASYNCIFY needed -
GpuDevice::readBuffer()is no longer documented as unconditionally blocking. Whether it blocks is a property of the backend: Metal, D3D11 and OpenGL read back in lockstep and fill the destination every call, while WebGPU cannot map synchronously and so trails the GPU by a frame or two. Callers must now own the destination across calls and treat a false return as "no new data yet" rather than an error - the previous contents stay valid -
ComputeParticlesDemo: keeps drawing the last particle snapshot on frames where no new one has landed, so it renders on the Emscripten WebGPU backend instead of showing nothing. The status label reports the landed-snapshot count alongside the frame count -
Component's effect path now reuses its offscreenGpuCanvasacross frames while the component size is unchanged, instead of allocating (and freeing) a full-size render target every frame. On a size change the outgoing canvas is released before the replacement is created, so itsRenderContextlease returns to the pool rather than forcing a second context to be reserved permanently -
ComponentEffectsDemo: shader effects now share a common base that compiles the pipeline at most once instead of retrying a failed compile on every frame, reports the compile error in the status label and on the console, and shows the CPU time spent applying the effect next to the paint time -
New
Layoutexample (examples/graphics): exercises the FlexBox and Grid layout containers with pages for direction, wrap, justify-content, align-items, align-content, flex grow/shrink/basis, align-self, order, margins/gaps, percentage sizing, min/max constraints, grid track sizing (px/fr/auto), explicit placement and spans, auto placement, grid alignment, and nested flex/grid compositions -
Added a
Toast Notificationsdemo to the graphics example app demonstrating theToastNotificationutility: a simplesendNotification, a richToastTemplate(attribution, actions, scenario, duration, expiration, event callbacks), an image template, and hide/clear -
Added a
Fluid Simulationdemo to the graphics example app: an incompressible Navier-Stokes solver (splat, vorticity confinement, divergence, pressure projection, gradient subtraction, semi-Lagrangian advection) ported from Pavel Dobryakov's WebGL Fluid Simulation, running every pass as a fullscreen-triangle GpuPipeline fragment shader over rgba8 ping-pong GpuTarget surfaces (sim fields packed as 16/24-bit fixed point) with a simplified bloom (original threshold/soft-knee prefilter into a small surface, blurred and upsampled with linear filtering, ordered dithering to hide 8-bit banding on fades) and shading composite presented viaGraphics::drawTexture— no CPU readback. Pipelines compile incrementally with a progress bar, and a calibration pass keeps sample orientation consistent across backends -
SDLComponentNativenow renders each window on its own dedicated render thread instead of the message thread: the component-tree walk runs under aMessageManagerLockwhile GL command submission and buffer swap happen unlocked, so multiple windows no longer serialize their frame rendering (and vsync waits) on the message thread -
Fixed
SDLComponentNative::repaint()calling-[NSWindow screen](viagetSize()→getWindowUnitsPerPoint()→SDL_GetDisplayForWindow()) from the render thread, which macOS's Main Thread Checker flags since AppKit requires that call on the main thread. It now uses the already up-to-datescreenBoundscached by the main-thread window event handlers instead of querying the display live -
Fixed
SDLComponentNative::runWithGraphicsContext()never invoking its callback on non-OpenGL desktop backends (Metal, Direct3D), silently dropping the work.Component::renderSubtreeOffscreen()routes through this hook, so any component with aComponentEffectset (e.g.ComponentEffectsDemo) rendered nothing on Metal/D3D;SDLComponentNative::runWithComputeContext()falls back to the same hook when no dedicated compute context exists, so GPU compute work initiated off the render thread was silently dropped there too -
Fixed Android crashing on any touch input after the render-thread rework: SDL invokes event-watch callbacks on the thread that pushes the event — the OS UI thread on Android, not the message thread — so the whole input/window-event pipeline mutated the component tree concurrently with the message and render threads. Events arriving off the message thread (window, input, drop and display events) are now marshaled onto it via
MessageManager::callAsync, with SDL-owned string payloads (SDL_EVENT_TEXT_INPUT,SDL_EVENT_DROP_FILE/DROP_TEXT) deep-copied since SDL frees them when the watch returns -
Mobile app lifecycle is now handled explicitly:
SDL_EVENT_WILL_ENTER_BACKGROUNDstops the render thread synchronously inside the event watch (the GL surface can be destroyed as soon as the callback returns, and SDL may block the message thread until the app resumes), andSDL_EVENT_DID_ENTER_FOREGROUNDrestarts rendering from the message thread.~SDLComponentNativealso removes its event watch before any teardown, so no other thread can dispatch events into a half-destroyed window -
Offscreen GPU work initiated from the message thread (component snapshots/effects via
runWithGraphicsContext()) now serializes against the render thread's frame submission on every backend:renderFrame()submitsGraphicsContext::end()/swap outside theMessageManagerLock, so Metal/D3D needed the same context lock the OpenGL path already used -
Fixed GPU resources leaking when released off the render thread: dropping a cached component canvas, snapshot, or GPU-backed
Imageon the message thread issued GL deletes with no context current (the render thread owns it now).GpuTargetandGpuTextureteardown routes through the newGpuDevice::runOnGraphicsContext(), which binds the rendering context on GL (mirroring the existingrunOnComputeContext()buffer-release path); the window destructor also keeps the GL context current while destroying its renderer and graphics context -
Fixed the window's GPU context activators dangling when GPU buffers or textures outlive their window: the activators now share a refcounted guard with
SDLComponentNative(the same lock that serializes GL context access), so a release after the window is gone locks, sees a null window and bails out instead of calling into freed memory -
Fixed the render thread overflowing its stack on complex paths: user
paint()code now runs on a secondary thread whose default stack is far smaller than the main thread's (512kB vs 8MB on macOS), and Rive'sGrTriangulatorrecurses hundreds of frames deep on many-point filled paths (e.g.SpectrumAnalyzerComponent's spectrum). The render thread is now created with an 8MB stack to match the main thread paint() previously ran on -
Fixed the render loop waiting forever when
framerateRedrawis above 250: the frame budget minus the 4ms pacing margin went negative, whichWaitableEvent::wait()treats as wait-without-timeout -
ComponentNative::setDesiredFrameRate()can now change a window's target framerate while it is rendering, and the newOptions::withUnfocusedFramerateRedraw()throttles a window to a lower rate whenever it does not have keyboard focus, for example while it sits behind a modal window. The unfocused rate only ever lowers the rate, andwithUpdateOnlyFocused()still takes precedence by stopping rendering entirely.getDesiredFrameRate()keeps reporting the rate that was asked for rather than the throttled one -
WebAssembly frames are now produced directly from
requestAnimationFramerather than from aTimer, removing a cross-thread handshake from the frame path: emscripten builds always pass-pthread, so the timer countdown is kept on the timer worker thread and reaches the main thread as a posted message that is drained on the next animation frame. The animation frame is also the only moment the browser presents, so a target rate below the display rate is now met by rendering every Nth callback instead of against a deadline, which can land just after a refresh boundary and cost a whole interval.SDLComponentNative::renderDrivenByTimeris removed in favour ofYUP_EMSCRIPTENifdefs. -
Added
Timer::getTimerFrequencyHz(), which reports the rate from the sub-millisecond interval instead of round-tripping it through whole milliseconds the way1000 / getTimerInterval()does.SpectrumAnalyzerComponent::getUpdateRate()now uses it and returns exactly the rate that was set -
Timernow keeps its period and countdown at sub-millisecond resolution and carries the overshoot of a late callback into the next period, so a rate that is not a whole number of milliseconds is honoured on average:startTimerHz (60)previously truncated1000 / 60to a 16ms interval and ran at 62.5Hz. The timer thread also measures elapsed time withTime::getMillisecondCounterHiRes()rather than the whole-millisecond counter.getTimerInterval()keeps its millisecond contract and now rounds, so a 60Hz timer reports 17 rather than 16, and a timer faster than 1kHz reports 1 rather than 0 -
Fixed timers never firing in a WebAssembly build without pthreads: the countdown is done by the timer thread, and the fallback that counted down on the message loop instead was behind
#if YUP_EMSCRIPTEN && ! defined(__EMSCRIPTEN_PTHREADS__), which never held for a target built withyup_add_standalone_app, the only place that flag is added. It is now selected at runtime by whether the timer thread is actually running, which also covers any other platform where it cannot be started -
The SDL render thread now paces frames on an absolute schedule instead of sleeping for whatever is left after each frame, so a frame's own overshoot is no longer carried into the next one, and it waits for a repaint request only until the point where
renderFrame()still fits inside the frame, using a rolling average of the measured render cost rather than a fixed 4ms margin (at 60Hz the old margin left 12.7ms of waiting, which plus any real render work overran the 16.7ms frame). On Apple the render thread is also started as a realtime thread, since outside the Mach time-constraint class the kernel coalesces timer expiries into a window of roughly 25% of the requested sleep -
Fixed glslang process initialization racing between concurrent
ShaderTranspilerconstructions (per-window render threads can now compile shader effects while a compute pipeline compiles elsewhere): the refcount was atomic but did not make a second thread wait forInitializeProcess()to complete; init/finalize now run under a lock. Also made the offscreen context pool'sleasedflag atomic in the GL and Metal devices, sinceRenderableTargetdestructors can run outside the context lock that serializes acquisition -
Fixed
WaitableTimer's non-Windows fallback overshooting frame deadlines by several milliseconds on macOS: it blocked for the entire remaining wait on a singlecondition_variable::wait_until, whose wake time is subject to OS scheduling / timer-coalescing latency. It now blocks for the bulk of the wait and closes the last few milliseconds with a tiered busy-wait againstTime::getMillisecondCounterHiRes(), restoring the precision the pre-WaitableTimerimplementation had -
SDLComponentNativenow caches its window units-per-point value:getWindowUnitsPerPoint()is only called when the window is created and on resize/dpi/screen-change events (SDL_EVENT_WINDOW_DISPLAY_CHANGED,DISPLAY_SCALE_CHANGED,PIXEL_SIZE_CHANGED,MOVED), and every other read (including from the render thread) uses the cached value, so no SDL window method is ever invoked from the render thread -
New
ComponentNative::vsyncflag /Options::withVSync()(off by default): synchronizes presentation to the display refresh (MetaldisplaySyncEnabled, GL swap interval, D3DPresent (1)). With vsync the present paces the render thread: after a frame that was presented and waited for the refresh, the next one starts without the software deadline, while frames that paint nothing (or backends whose present doesn't wait, like Dawn) keep it so the loop can't spin
- Rive runtime bumped from v0.1.62 to v0.1.155
- GraphicsContext GPU context integration. New
GraphicsContext::isGpuAvailable()capability probe;gpuContext()is retained but documented@internalas the single backend bridge. - New
GpuTextureclass (rhi/yup_GpuTexture.h): opaque reference-counted GPU texture wrappingrive::gpu::Textureorrive::gpu::RenderCanvas. Obtained fromGpuCanvas::asTexture()or constructed internally byImage::fromTexture(). - New
GpuTargetclass (rhi/yup_GpuTarget.h): low-level render-pass-only offscreen GPU surface (create,beginRenderPass,asTexture,asImage,readPixels). Its backing texture is allocated from the context's main render context, so it does not reserve a dedicatedrive::gpu::RenderContext- use it for customGpuPipelinework (e.g. post-process passes) that needs no 2D drawing. - New
GpuCanvasclass (rhi/yup_GpuCanvas.h): consolidated backend-agnostic offscreen GPU surface that now composes aGpuTarget(over aRenderableTarget) and creates a non-owningGraphicslazily only when 2D drawing is requested. - Python bindings now expose
GpuColorasyup.GpuColor(backingGpuRenderOptions.clearColor), comparable withyup.Color. - Python bindings now expose
GpuLoadOp/GpuStoreOpandGpuRenderOptions.loadOp/.storeOp.GpuRenderOptions.clearis kept as a bool view ofloadOp, soGpuRenderOptions(True, color)andopts.clearstill read the same as before.
- New
yup_rhimodule: the GPU abstraction layer extracted fromyup_graphicsinto its own module (depends onyup_core,yup_shading,rive_renderer). All RHI classes (GpuFrame,GpuPipeline,GpuBuffer,GpuTexture,GpuTarget,GpuRenderPass,GpuPipelineCache) now live inyup_rhi.yup_graphicsdepends onyup_rhifor GPU access. GpuDevice: new reference-counted GPU device abstraction (wasGpuContext). Owns the native GPU device and command queue without requiring a window - can be used for headless GPU compute (e.g. audio DSP on the GPU). Created viaGpuDevice::create(GpuPlatform, Options). All RHI factory methods (GpuFrame::begin,GpuPipeline::compile*,GpuBuffer::create,GpuTarget::create) now takeGpuDevice::Ptrfor safe shared ownership.GpuPlatformenum: standalone platform enum (Headless,Metal,Direct3D,OpenGL,OpenGLES,WebGPU) replacing the nestedGpuDevice::Api.GraphicsContext::getPlatform()returns it directly - no typedef alias.GpuColorstruct (rhi/yup_GpuTypes.h): lightweight 4-component GPU color for render options. Implicitly constructable from any type withgetRedFloat()/getGreenFloat()/getBlueFloat()/getAlphaFloat()(e.g.yup::Color), soGpuRenderOptions { true, Colors::transparentBlack }works without code changes.GraphicsContextsimplified: wraps aGpuDevice::Ptr(obtained viagetGpuDevice()returningGpuDevice::Ptr). Offscreen target management (createOffscreenTarget,beginOffscreen,endOffscreen,readOffscreenPixels) delegated toGpuDevice. Factory accepts optionalGpuDevice::Ptrto share an existing GPU device.- Backends:
GpuDevicehas native implementations for all platforms (Metal, OpenGL, Direct3D 11, Dawn, WebGPU/Emscripten, Headless). OpenGL backend probesGL_VERSIONat runtime to detect compute shader support (GL ≥4.3 / GLES ≥3.1). ::Ptrsafety: all RHI types that own resources (GpuPipeline,GpuBuffer,GpuTexture,GpuTarget,GpuCanvas) are reference-counted with::Ptr. Factory methods takeGpuDevice::Ptrto keep the device alive for the resource's lifetime.GpuFrameis move-only stack RAII and takesGpuDevice&(no ownership).
- New
GpuComputePipelineclass (rhi/yup_GpuComputePipeline.h): an immutable compiled compute pipeline that bypasses ore to go directly to the backend-native API (MetalMTLComputePipelineState, D3D11ID3D11ComputeShader, WebGPU/Dawnwgpu::ComputePipeline, OpenGLGL_COMPUTE_SHADER).compile(ctx, source, GpuWorkgroupSize),compileFromBundle(ctx, ShaderBundle), andcompileFromGlsl(ctx, glsl)(whenYUP_ENABLE_SHADER_TRANSPILER = 1) all returnResultValue<GpuComputePipeline::Ptr>. - New
GpuComputePassclass (rhi/yup_GpuComputePass.h): move-only RAII compute dispatch encoder (GpuComputePass::begin(device)). Binds aGpuComputePipeline, storage buffers (setStorageBuffer), uniform buffers (setUniformBuffer), and textures (setTexture), then dispatches workgroups viadispatch(gx, gy, gz). GpuBufferextended withGpuBufferType::storage: native storage buffer creation for each backend (MetalMTLBuffer, D3D11 structured buffer + UAV, WebGPUStoragebuffer, OpenGLGL_SHADER_STORAGE_BUFFER). Storage buffers are bound toGpuComputePass::setStorageBuffer().GpuDevicebackends expose native compute handles:getDevice()/getCommandQueue()(Metal),getD3DDevice()/getD3DDeviceContext()(D3D11),getWgpuDevice()/getWgpuQueue()(WebGPU/Emscripten),getBackendDevice()/getDevice()/getQueue()(Dawn).GpuAudioProcessingDemoexample: real-time GPU-accelerated audio effect (gain + soft clipper) using compute shaders. Captures live audio viaAudioIODeviceCallback, uploads to GPU storage buffers, dispatches a compute shader, and reads back processed audio - all on the audio I/O thread.- New
GpuDevice::updateBuffer(): writes new data into an existing storage buffer without reallocating it (Metalcontentsmemcpy, D3D11UpdateSubresource, WebGPU/DawnWriteBuffer, GLglBufferSubData). FixesGpuAudioProcessingDemoreallocating its input storage buffer every audio callback, which caused audible stutter. The gain/mix parameters remain a uniform buffer (as before) - that path is unaffected and its small per-dispatch allocation is negligible next to the audio-block-sized buffer this fix removes. - Fixed
ShaderTranspiler's MSL backend assigning storage/uniform buffer indices via spirv-cross's own auto-incrementing scheme instead of the shader's declaredlayout(binding=N): addedCompilerMSL::Options::enable_decoration_binding = trueso the compiled[[buffer(N)]]index always matches the declared binding, matching whatGpuComputePass's native dispatch (which binds slots asgroup*16+bindingwith no reflection indirection) requires. This was silently producing zero output from any Metal compute shader with more than one storage/uniform buffer, includingGpuAudioProcessingDemo. - Metal
GpuDevice/GpuComputePasscalls now wrap their Objective-C work in@autoreleasepoolblocks - without one, real-time callers (e.g. an audio thread with no ambient pool) accumulated command buffers/encoders indefinitely. - Metal
GpuDevice/GpuComputePasscalls now wrap their Objective-C work in@autoreleasepoolblocks — without one, real-time callers (e.g. an audio thread with no ambient pool) accumulated command buffers/encoders indefinitely. - Fixed
ShaderTranspileremitting ESSL 3.00 for every stage, which made spirv-cross reject compute shaders on OpenGL ES ("At least ESSL 3.10 required for compute shaders") — compute stages now target ESSL 3.10 (#version 310 es) while vertex/fragment stages keep ESSL 3.00.
- Added TIFF read/write support (
TiffImageFormat) via libtiff: RGB, RGBA, Grayscale at 8/16-bit; multi-page reading; DPI and EXIF/ICC/XMP metadata extraction. - Added TGA read/write support (
TgaImageFormat): uncompressed and RLE-compressed truecolor and grayscale variants; RGB and RGBA output with alpha channel preservation. - Added animated WebP encoding and decoding support to
WebPImageFormatWriterandWebPImageFormatReader: per-frame metadata (canvas dimensions, frame count, loop count, per-frame delays, dispose/blend modes) and frame decompression with manual compositing. - Added animated PNG (APNG) encoding and decoding support to
PngImageFormatWriterandPngImageFormatReader: manual chunk-level parsing ofacTL/fcTL/fdATchunks for animation metadata, per-frame libpng decoding via synthetic minimal PNG construction, and canvas compositing supporting all three APNG disposal operations (none, background, previous) and both blend operations (source, over). ImageFormat::Optionsstruct controls metadata extraction:.withMetadata(true)enables text metadata and DPI;.withRawChunks(true)enables raw binary chunks (EXIF, ICC, XMP). When both are false (the default),ImageMetadatais not allocated - true zero overhead.- Introduced a ref-counted
ImageMetadataobject (ImageMetadata::Ptr) attached toImageandImageFormatReader::metadata. DPI, text entries, and raw binary chunks are all accessed through the metadata object only when requested viaOptions. - Lossless roundtrip tests for all formats (BMP, PNG, WebP, TGA, TIFF, PPM, GIF) now verify pixel-perfect fidelity after write→read; animated roundtrip tests for GIF, WebP, and PNG verify per-frame pixel integrity.
StyledText::TextModifier::appendText()gained aColoroverload that creates (and caches per color) a solid fill paint, andGraphics::fillFittedText()now honors per-run style paints when every run carries one - enabling syntax-colored text. Single-colorStyledTextusage is unchanged.FontgainedisEmpty().- New
Fontstatic loaders:Font::loadFontFromData(),Font::loadFontFromFile(),Font::loadFontFromFirstAvailableFile(),Font::loadSerifSystemTextFont()andFont::loadMonospaceSystemTextFont(), all returningResultValue<Font>(wasOk()/failed()/getValue()). The former theme-local system font lookup helpers moved intoFont; macOS/iOS use the CoreText system UI fonts, other platforms try well-known system font files. ImageFormatReaderandImageFormatWritergained adeleteSourceWhenDestroyedconstructor parameter, defaulting totrueso existing callers are unaffected. It lets a format be handed a stream it did not open and either take ownership of it - deleting it on destruction, as the C++ path has always done - or decline and leave the caller to own it, in which case the reader's / writer's owning pointer is released rather than deleted. The reader overload that accepts the caller'sInputStreamalongside the flag is what letsImageFormatManager::createReaderForhand a Python format the stream it opened without leaking it
-
Input follows what is displayed:
ComponentEffect::displayToContent()/contentToDisplay()let distorting effects route pointer input, and the newComponent::getChildPointFromLocal()/getLocalPointFromChild()hooks, used by every input path, let a parent present children through a custom projection.Component::setManuallyComposited()andrenderToTexture()let a parent render a live child subtree itself, for example on a 3D surface. When painted content moves under a stationary pointer, the SDL window now re-evaluates the pointer after the frame and sends the enter/exit, move or drag events it implies.ComboBoxpopups open inside the closest transformed or manually composited ancestor (Component::getPopupParentComponent()), so they are presented with the widget. -
Fixed a component's opacity being applied twice when it has an effect or is cached to texture: the offscreen content is rendered opaque and the opacity is applied once when compositing.
-
Fixed
Component::getTransformToScreen()adding a desktop window's position twice and applying the root's own transform; it now matcheslocalToScreen(). -
Component::findComponentAtForMouseEvent()maps the point into children through transforms, effects andgetChildPointFromLocal(), likefindComponentAt(), so mouse input reaches transformed, warped and manually composited components. -
Effect and cached-to-texture canvases are rendered at the display scale, so they stay sharp on high-density displays.
Component::renderToTexture()takes ascale(passg.getContextScale()). Snapshots stay at one pixel per point. -
Fixed animation repaint requests bypassing the render deadline, which produced uneven frame submission and could run the Metal render loop uncapped when VSync was enabled. Repaint events now wake the render thread without allowing a frame before its scheduled deadline, and missed deadlines skip directly to the next future frame
-
Animation precomposition canvases are now leased from a bounded per-frame scratch pool. Their cache previously used the sampled animation phase as a persistent key, allocating another full-size GPU canvas for each new animation frame and retaining all of them until the renderer was reset
-
Componentgained the callbacks it was missing for reacting to its surroundings:childBoundsChanged (child)on a parent whose direct child moved or was resized (a singlesetBounds()reports once, not once for the move and once for the resize),parentSizeChanged()on each direct child after its parent'sresized()has run,indexInParentChildrenChanged (oldIndex, newIndex)on a component whose z-order position changed, andfocusOfChildComponentChanged (child, cause)on every ancestor of a component that gained or lost the keyboard focus -
Component::keyStateChanged (key, isDown)reports a key going down or coming up even when the press never reacheskeyDown(), and does not retrigger on auto-repeat, which makes it the place to track a held modifier or chord key.Component::modifierKeysChanged (modifiers)reports the modifier state changing, including when it changes during a mouse gesture rather than a key event -
Component::hitTest (x, y)decides whether a local point counts as being inside a component, and overriding it overrides hit testing for the mouse: a non-rectangular widget can return false for its transparent corners and let events through to whatever is behind it. An override can only take area away, since a point outside a component's bounds is rejected by its parent beforehitTestis consulted -
Fixed mouse ups going missing. The windowing backend polled the live OS button state on a timer and dropped the whole gesture whenever it saw no button down, but that state changes the moment the button is physically released, while the matching
SDL_EVENT_MOUSE_BUTTON_UPis still queued: a tick landing in between left the following mouse up with no component to deliver it to, so the component stayed pressed and never saw the click. The poll is gone; a gesture whose release was genuinely consumed elsewhere - a native drag session runs its own event loop and swallows it - is now ended explicitly through the newComponentNative::cancelCurrentMouseGesture(), which delivers the missing mouse up rather than silently forgetting it. -
Fixed a window losing mouse capture after a drag.
setGlobalMouseCaptureActive()and the window's owncaptureMouseshared a single flag to track a reference on one process-wide capture count, so a drag starting from a window that already captured released that window's capture when it ended, and asetVisible()during a drag released the drag's. Each now owns its own reference. -
Drag and drop can now leave the application.
DragOptions::withExternalDragAllowed (true)opts a source in, and the manager hands the gesture to the platform once the pointer is over none of our windows: on macOS that runs a native AppKit drag session carrying files, text and PNG images, so a YUP drag can be dropped on another application. An image payload is already PNG encoded byDragAndDropData, so it goes onto the pasteboard as it is. Windows and X11 have no implementation yet and keep such a gesture in the app - and note that a source cannot be both in-app and external, because crossing between two windows is indistinguishable from leaving the application. -
New "Drag and Drop" example in
examples/graphics: two trays of tiles that drag between each other and into a second window opened from the demo (live tiles are registered process-wide, so a tile can move across windows), a drop zone that accepts files and text dragged in from another application and lists what arrived, and a list whose rows drag out with multiple selection - starting a drag on one of several selected rows carries the whole selection. -
ListBoxis now a drag source: dragging a row asks the model forgetDragSourceDescription (selectedRows)and, when that returns anything other than a default-constructedvar, begins a drag carrying it. A string description is mirrored into thetextMIME type as well, so a target that reads only MIME data still sees something. NewcreateDragSourceComponent (selectedRows)produces the drag image and can be overridden; the default is a circle carrying the number of dragged rows.setDragSourceEnabled (false)makes a list undraggable without consulting its model at all. -
Fixed
ListBoxmultiple-selection behaviour: a plain click now replaces the selection, having previously added to it becauseselectRow()only ever adds in multiple selection mode. Shift extends a range from the last click without shift and keeps that anchor across further shift-clicks, so repeated shift-clicks grow or shrink one range instead of creeping along a row at a time, and command/control toggles an item. The range branch also repaints now, which it never did - the rows really were selected, but nothing showed it. -
Fixed modifiers being misread on every mouse and touch event: the windowing backend built a
KeyModifiersstraight from SDL'sSDL_GetModState()bitmask rather than through the mapping the keyboard path uses. SDL's bit layout does not overlapKeyModifiers's masks, so control, command and alt never registered at all, and shift only appeared to work because SDL's left-shift happens to sit on the same bit asshiftMask. -
Fixed translucent windows never being cleared, and the compositor discarding their alpha anyway. The renderer only used
LoadAction::clearfor a window that renders continuously, so a window that draws on demand kept whatever the drawable held before; it now clears whenever the clear colour is not opaque. Alongside that, the Metal layer no longer forces opaque, a Windows swap chain created for a composed window requestsDXGI_ALPHA_MODE_PREMULTIPLIED, and desktop OpenGL asks for an alpha channel as GLES already did. A window created withComponentNative::Options::withTransparent (true)now composites properly rather than showing its un-cleared corners. -
Drag and drop is now opt-in mixins plus an app-global session rather than part of
Component(see Breaking changes).DragAndDropSourcestarts a drag from a payload and an optional ghost,DragImageComponentis the borderless window that renders that ghost, andDragAndDropManagerowns the drag in flight: it listens for global mouse events for its duration, finds the component under the cursor across every native window (Desktop::findComponentAt), and drives the enter/move/exit and drop dispatchers. The performed action defaults to copy, or move with Shift held.DragAndDropTargetComponentis a convenience for a component that is also a target. Drags that pass over a window from outside the application arrive through the same machinery; note that the OS reports no payload until the drop itself, so a target that wants to react while such a drag hovers must accept an empty payload. -
New
ComponentNativewindow optionswithTransparent(),withAlwaysOnTop()andwithFocusable(), mapping onto SDL's transparent, always-on-top and not-focusable window flags, plusComponentNative::getComponent()andsetGlobalMouseCaptureActive()so a window can keep receiving mouse input while the pointer is outside it. -
Fixed
Component::addChildComponent()(and thereforeaddAndMakeVisible()) not detaching a component from its previous parent. Reparenting left the component in the old parent's child list as well, so it was still painted and laid out there whilegetParentComponent()pointed at the new one, and a further reparent added it a second time. The old parent is now told before the new one adopts it. -
Fixed
SDLComponentNative's destructor resurrecting theDesktopsingleton during shutdown: unregistering a window calledDesktop::getInstance(), which constructs a new desktop when teardown has already destroyed the old one, tripping the check at the end ofDeletedAtShutdown::deleteAll(). It now unregisters throughgetInstanceWithoutCreating()and skips it when the desktop is already gone. -
DeletedAtShutdownnow asserts at the point a new object is created whiledeleteAll()is deleting the others, instead of failing anonymously once the pass has finished and the stack naming the culprit has been unwound. -
New
ComponentNative::RepaintMode(ComponentNative::Options::withRepaintMode) selects how a window turns its accumulated dirty rectangles into repaint work. The default,RepaintMode::disjointRegions, repaints each dirty rectangle in isolation: a parent shared by several dirty rectangles is painted once, clipped to those rectangles, so components lying between two distant dirty rectangles are no longer repainted.RepaintMode::boundingBoxkeeps the previous behaviour of collapsing every dirty rectangle into one bounding box (repainting everything in between) and remains available as a fallback. -
Transformed components are no longer clipped to their untransformed bounds.
Component::internalPaint()now maps its local bounds through the transform accumulated from the component and its ancestors to the top level component before intersecting them with the accumulated dirty rectangles, so the redraw area of a scaled, rotated, or sheared component matches the area it actually covers and parts of it falling outside the untransformed bounds are painted again. -
A
paint()that throws no longer leaves the graphics frame open.SDLComponentNative::renderFrame()calledcontext->begin()inside its render lambda butcontext->end()only after that lambda returned, so an exception from user paint code skipped theend()/tick()pair and the next frame began on a context that had never been flushed. The pair now runs from a scope guard -
GUI: multitouch input on any platform with touch hardware (mobile, Emscripten in a mobile browser, and desktop touchscreens). Every finger is delivered through the existing mouse callbacks (
mouseDown/mouseDrag/mouseUp, left button held): the first finger behaves exactly like a mouse, each additional finger arrives in parallel with a stable, dense finger index exposed asMouseEvent::isTouch()/MouseEvent::getTouchIndex(), plus touch pressure asMouseEvent::getPressure()(0.0-1.0). Each finger is hit-tested independently, keeps its index for the whole contact, and has its own double-click detection. SDL's synthetic touch-to-mouse events are disabled so every touch is delivered exactly once. -
New "Touch Trails" example in
examples/graphics: draws one hue-spaced colored trail per pointer (each finger gets its own color viaMouseEvent::getTouchIndex(), with a mouse as a single pointer on desktop). Points carry an age in fade-timer ticks - not the wall clock - and their size and opacity decay with that age, so late frames can't make trails jump or flicker; the fade timer also runs while a pointer is down, dissolving the trail until only the pressed point remains, and once the finger lifts that point fades out too. -
ApplicationThemenow exposessetDefaultMonospaceFont()/getDefaultMonospaceFont(), mirroring the existing default and icon font APIs. The default theme populates it from the embedded (or system) monospace font. -
The default theme now embeds JetBrains Mono Variable (SIL OFL) as its monospace font when
YUP_EMBED_DEFAULT_THEME_TEXT_MONOSPACE_FONT = 1(forced on Emscripten), falling back to the system monospace font otherwise. Thetools/embed_font.pyregenerates the.incbyte arrays from any font file. -
New
CodeDocument(line-based text model withUndoManager-backed edits, positions, and incremental change notifications),SyntaxDefinition(JSON-driven language descriptions loaded from data/files or the built-in C++ / GLSL / Python definitions),CodeTokeniser(incremental per-line tokenizer with a line-state machine for multi-line constructs and lazy re-tokenization), and theCodeEditorcomponent: syntax-highlighted editing, caret/selection with anchor semantics, clipboard, undo/redo, read-only, smart auto-indent, an optional line-number gutter with breakpoint markers, find/replace (find-all, next/previous with wrap, replace-one, replace-all in one undo step, match highlighting), bracket matching, and an optional minimap overview. Defaults to the theme's monospace font. Seedocs/ui/code-editor.md. -
Added a built-in XML
SyntaxDefinition(available asSyntaxDefinition::getBuiltIn ("xml")and matched for.xml,.svg,.html,.xamland other markup extensions), with<!-- -->block comments, tag/attribute punctuation and<? ?>/<!/<///>operator highlighting. TheCodeEditordemo now has a language dropdown to switch between the built-in C++ / GLSL / Python / XML definitions. -
Added a built-in YDSP
SyntaxDefinition(SyntaxDefinition::getBuiltIn ("ydsp"), matched for.ydsp), covering the language's keywords, primitive types and Faust-style composition operators (<:,:>,->,~, …). Used by theyup_dsp_jit"YDSP Synths" example's new Editor tab. -
Fixed
CodeDocument:newLineCharswas default-constructed to an empty string instead of"\n", breakinggetText(),getTextInRange(), and character-offset calculations for all multi-line documents;applyEdit()returned a wrong caret column for single-line insertions (omittedstartIndex), making every subsequent undo call operate on an inverted range and silently no-op; removed theendsWithNewlinespecial case that returned a pre-newline position and similarly broke undo for Enter at the beginning of a line or in the middle of a line. -
Fixed
CodeEditor:undo()andredo()now clampcaretPositionto the new document length and clear the selection after each operation, preventing an out-of-bounds caret after undo shrinks the document;replaceNext()now uses the position returned byreplaceRangeinstead ofselectionStart + replacement.length(). -
Fixed
CodeTokeniser: cutting or deleting text that removes one or more lines left the token cache larger than the document and the forward-propagation stability check could declare a line whose content had shifted "unchanged", returning stale tokens (wrong syntax colors) for every line below the cut point.codeDocumentChangednow shrinks the cache to the new document line count and proactively marks all shifted lines dirty before the stability pass runs. The same problem existed in the other direction and was more visible in practice: inserting a line (pressing Enter, or a multi-line paste) grew the document butcodeDocumentChangedhad no branch for it at all, so every cached entry at or after the edit point kept referring to whatever used to be at that index - one or more lines off from where it actually was - and the stability check could decide a shifted-in line's state was "unchanged" and never mark it dirty, leaving it with stale, wrongly-sized tokens that fail to tile the line and fall back to unhighlighted plain text. Both directions are now handled the same way: resize to the new line count and mark everything from the edit point to the new end dirty, so misaligned cache entries are always discarded and recomputed from the live document text rather than reused. -
Fixed
CodeDocument::setText()freezing for several seconds on a large paste: its line-splitting helper indexed the (UTF-8-backed) inputStringby character position inside the split loop, and bothoperator[]andlength()are O(n) for UTF-8, turning the split into O(n²). It now walks the text once with aCharPointer. -
Fixed
CodeEditordrawing selected/highlighted/caret text past the gutter and minimap when scrolled horizontally, since nothing clipped that content to the text area; the gutter's painted background also stopped 4px short of where the text area actually starts (disagreeing with the hit-test boundary used bymouseDown()), reading as misalignment on the left. The minimap overview now merges lines that map to less than one device pixel row into a single bar instead of issuing onefillRect()per source line on every paint regardless of visibility. -
Fixed
StyledText::update()callingFont::getPath()(a CoreText round-trip on Apple platforms) once per glyph occurrence instead of once per unique glyph; a glyph's outline is the same every time for a given font, so it's now cached and reused, cutting a measured 240ms of the 453ms spent reshaping text on a single keystroke. -
Fixed
CodeEditorreshaping (tokenizing, laying out and re-tessellating) the entire document on every single edit, making typing in a large file cost seconds per keystroke (measured: 2s for one backspace in a large file, mostlyStyledText::update()).styledTextnow only ever holds the currently visible lines rather than the whole document; scrolling reshapes just the newly-visible range. Selection, search highlights, the caret, and Up/Down arrow navigation were adjusted to work correctly when their target is outside the currently-shaped range (falling back to an exact document-position computation rather than depending onstyledText). Components that never callsetSize()/setBounds()on theirCodeEditor(as none of its unit tests do) keep shaping the whole document, since there's no meaningful "visible range" to restrict to without a real size. -
CodeEditornow renders through aCodeEditorScheme(newcode/yup_CodeEditorScheme.h): every color — background, gutter, caret, current line, selection, search highlight, breakpoint and the per-token syntax colors — is stored keyed byIdentifier(CodeEditorScheme::setColor/getColor, string constants inCodeEditorScheme::ColorId) and switched withCodeEditor::setScheme. Built-in well-known schemes are provided viaCodeEditorScheme::getBuiltIn:monokai,alabaster,oneDark,solarizedDarkandsolarizedLight. The editor's painting moved into the theme (Themes v1) as a registeredCodeEditorcomponent style, and a vertical auto-hideScrollBarnow appears when the document overflows the viewport. TheCodeEditordemo gained a scheme dropdown. -
Added responsive Rive Layout support to
Artboard: per-node listeners (setNodeBoundsListener/clearNodeBoundsListener/clearAllNodeBoundsListeners/getNodeBounds) and attached components driven by layout reflows (attachComponentToNode/detachComponentFromNode/detachAllComponents, withNodeAttachmentOptionscontrolling whether the component fills the node's bounds or only tracks its position, follows its rotation viaComponent::setTransform, and — in track position mode — which component pivot point (pivot) is anchored to which point of the node's bounds (anchor), each accepting anyJustification). Node listeners fire whenever the node's bounds or its on-screen orientation (rotation, scale, skew, mirroring) change and hand the callback a cachedArtboardNodehandle — the same objectArtboard::findNodereturns — instead of allocating one per event, so the bounds/transform are read from the handle (which now also exposesgetViewTransform(), the node's transform in the same component coordinates asgetBounds()) and handles invalidate safely when the artboard is cleared or its file replaced. Reflows require the artboard's own layout (setFitting (std::nullopt)) orFitting::fill, which propagate the component size into the Rive artboard (leaving those fit modes restores the artboard's authored size, and non-layout nodes report their real shape geometry instead of a unit rect). -
Artboard::setFile()now takes an optional artboard name, loading that named artboard from the file (viarive::File::artboardNamed) instead of always loading the file's default artboard. -
Added
tools/rive_inspect.py, a standalone Rive (.riv) binary inspector withnames,info, andtreecommands (list named objects, dump object properties, and print the artboard hierarchy), resolving keys against the vendored rive generated headers without needing the runtime. -
Added Rive ViewModel (data binding) support to the artboard module:
ArtboardFileexposes the ViewModel schemas stored in a .riv file (getNumViewModels,getViewModelNames,getArtboardViewModel,createArtboardViewModelInstance) through new refcountedArtboardViewModel(schema introspection: typedPropertyInfoper property with input/output flags and enum options, authored-instance names) andArtboardViewModelInstancehandles (typed and genericvarproperty get/set by name or dotted path into nested viewmodels and lists, enum selection, trigger firing, list add/remove/swap/clear, and an optionalsetPropertyChangedCallbacknotified synchronously on value changes, including output bindings applied by the artboard).ArtboardgainedbindViewModelInstance/unbindViewModelInstance/getBoundViewModelInstance/getViewModelName, so a file's default instance can drive the artboard's data-bound properties and state machine transitions; a bound instance must come from the sameArtboardFile. -
Artboard::setAllInputs()is now implemented (it was a documented no-op): it applies anArray<var>snapshot produced bygetAllInputs()back throughsetInput(), matching entries by their"id"and ignoring unknown ones, so a snapshot can be saved and restored or moved between artboards. Triggers are stateless and deliberately carry no"value"in either direction. -
ArtboardViewModelInstancepaths may now terminate on a list index ("items.2"), which names the item's viewmodel instance and resolves throughhasProperty()andgetNestedInstance(); previously only value paths through an index ("items.2.quantity") worked and the two-segment form silently reported nothing. -
ArtboardViewModelInstancestructural list changes (addListItem,addListItemAt,removeListItem,swapListItems,clearListItems) now notifysetPropertyChangedCallbackwith the list's own path and an emptyvar, as the callback's documentation already promised. Appending an item now extends the observer tree instead of rebuilding it wholesale, so filling a list is no longer O(N²) in observer allocations. -
Artboardnow auto-detaches a component attached withattachComponentToNode()when that component is destroyed, and attaching a component that already follows another node moves it rather than leaving it driven by both. -
Artboardnow re-derives an attached component's position when the component is resized. IntrackPositionmode the component keeps its own size and that size is what thepivotis measured against, so resizing it after attaching used to leave the position stale until the node itself moved — which for a static layout is never. Size can now be set before or afterattachComponentToNode()with the same result. InfillNodemode the node owns the size, so an owner-driven resize is put back. -
Added
docs/ui/artboard.md, covering file loading and asset resolution, layout/alignment, state machine inputs and events, node access and component attachment, and the ViewModel data-binding flow. -
Artboard::getViewModelName(),ArtboardFile::getViewModelNames()andArtboardViewModelInstance::hasProperty()are nowconst/ no longernoexceptwhere they allocate;ArtboardViewModelandArtboardViewModelInstanceno longer type-erase their Rive pointers throughvoid*, andArtboard/ArtboardFilegained the leak detector and non-copyable declarations the rest of the module already carried. -
Fixed
Component::getScreenPosition()andMouseEvent::getScreenPosition()adding the source component's own offset twice.getScreenPosition()computedlocalToScreen (getPosition()), butlocalToScreen()already addsgetPosition(), so any nested component reported a screen position shifted by its own parent-local offset (andMouseEvent::getScreenPosition()inherited the same fault). Both now map throughlocalToScreen()exactly once, sogetScreenPosition()agrees with the origin ofgetScreenBounds(). -
GUI: drag-and-drop payloads can now carry arbitrary MIME data and a same-process native object.
DragAndDropDatastores anArray<ClipboardData>plus avarand offerswithImage/getImage(PNG-encoded), the genericwithMimeData/getMimeData/getMimeTypes/getAllMimeData, andwithNativeObject/getNativeObject, with files and URIs sharing thetext/uri-listMIME type. The class lives inmodules/yup_gui/dragdrop/. -
GUI: drag-and-drop targets are now an opt-in
DragAndDropTargetmixin instead of fiveComponentvirtuals, soComponentstays free of drag-and-drop and the whole feature is isolated inmodules/yup_gui/dragdrop/. The SDL backend resolves targets with adynamic_castand bubbles from the deepest component under the cursor, preserving the previous enter/move/exit and drop semantics;DragAndDropSourceDetailscarries the payload, source component, target-local position, allowed actions and suggested action. Seedocs/ui/component-drag-and-drop.md. -
GUI: added
DragAndDropTargetComponent, a concreteComponentthat is also aDragAndDropTarget, for the places that need one nameable type (factories, containers, the language bindings). The Python bindings are complete again:yup.DragAndDropDatagainedwithImage/getImage/hasImage,withMimeData/getMimeData/getMimeTypes/hasMimeDataandwithNativeObject/getNativeObject/hasNativeObject, andyup.DragAndDropAction/yup.DragAndDropActions,yup.DragAndDropSourceDetailsandyup.DragAndDropTargetComponentare bound - a Python subclass of the last overrides the target callbacks (or assigns theonItem*callables) and now actually receives drops, because it subclasses a single C++ type rather than deriving fromComponentandDragAndDropTargetseparately (which pybind would give two unrelated C++ subobjects that thedynamic_castdispatch could never find).
SpectrogramComponentnow keeps its waterfall history on the GPU: a precompiled.yslshader bundle (embedded inyup_SpectrogramComponentShader.inc, built with theyup_shader_bundlerhost tool) drives a single fullscreen-triangleGpuRenderPass(seeGpuPipeline) that scrolls the previous frame down by the pending rows and writes the new rows with the color map applied entirely on the GPU, uploading only the raw magnitudes as a uniform buffer - no per-paint CPU pixel upload, no GPU texture creation, and no 2D canvas flush (if the bundle cannot be compiled no waterfall is rendered). Pending FFT rows are always consumed (applied or dropped) so the update queue can never accumulate. The log-frequency → FFT-bin mapping is precomputed once per configuration instead of recomputed with pow/log per row, and the frequency grid (lines + labels) is cached in an offscreen canvas and only re-rendered when the frequency range or size changes. The component now requires a GPU render context (the CPUImagefallback was removed). The component's per-framerefreshDisplayhook processes pending FFT rows, and the history is presented at a fractional vertical offset that slides the newest row into place over one row period; that offset is pulled back with a fixed time constant instead of being clamped, so the motion never stalls at the edge of its window. The scroll speed is adjustable via the newsetScrollSpeed()multiplier (1.0 = realtime, 0.0 = paused).SpectrogramComponent's waterfall scroll is now independent of the frame rate. The component requests a repaint on every frame while the waterfall is live (through the existingrefreshDisplayhook) instead of only when an FFT row arrives, so the sub-row offset above is actually rendered - previously the repaint cadence matched the row arrival cadence exactly, which made the display advance one whole row per repaint (one pixel, for a component whose height matches the history). A frame that produced more FFT rows than one GPU pass can write (a pass writesdefaultSpectrogramMagnitudes / defaultSpectrogramWidthrows) now drains them over several passes instead of dropping the surplus, which is what happened whenever the display ran below the FFT row rate: at 30 fps half the history of a 2048/1024 configuration was silently lost.setScrollSpeed (0.0)now truly freezes the waterfall (pending rows are discarded rather than written, so the content no longer keeps sliding down while paused), and the animation stops requesting repaints once the analysis stalls instead of spinning on a frozen frameSpectrogramComponentwaterfall failures (shader bundle load, pipeline compile, and GPU pass encode/draw) are now reported viaLogger::outputDebugStringin all build configurations instead of silently dropping pending rows, and the waterfall texture's render resolution is exposed as the newdefaultSpectrogramRenderWidthconstant (2x the frequency-bin count -getSpectrogramImage()returns that full-resolution image).- Fixed
SpectrumAnalyzerStatenever flagging FFT data as ready after a single bulkpushSamples(): the readiness check ran before the scoped FIFO write had committed (theAbstractFifo::ScopedWritecommits in its destructor), soisFFTDataReady()stayed false until a second push arrived.pushSample()/pushSamples()now commit the write before checking, so a pushed window is immediately available toSpectrogramComponent::refreshDisplay()instead of leaving the backlog untouched. SpectrumAnalyzerComponentandSpectrogramComponentno longer snap every display point to its nearest FFT bin, which rendered identical levels (a flat staircase) for all the points sharing one bin. The newSpectrumBinMappinghelper (displays/yup_SpectrumBinMapping.h) holds the shared log-frequency → fractional FFT bin mapping and evaluates levels continuously: the three bins surrounding a fractional position are parabolically interpolated in the amplitude decibel domain, evaluated at that position rather than at the vertex of the parabola, with a monotone linear fallback where the neighbours are not concave so a steep bin pair cannot undershoot. Display bands are aggregated over their fractional edges -peakfor the peak/RMS level modes,sum(the levels integrated across the band width, in bin units) forpowerDecibelsandmeanforpowerSpectralDensity- so bins entering or leaving a band no longer step the displayed level. In the power modes this makes the band power an integral over the band's bandwidth instead of a sum over the integer bins its edges happen to touchSpectrumAnalyzerComponentbuilds its spectrum outline once per pixel column, interpolating between the 512 smoothed display points, so the curve stays continuous at any component width and at HiDPI scales instead of following the 512-point polyline verbatim. The privatecomputeSpectrumPath (Path, …)becamecreateSpectrumPath (const Rectangle<float>&, bool)returning aPath, removing its reliance onPathsharing itsrive::rcp<RiveRenderPath>between copies
The FlexBox and Grid containers landed in this cycle (they were previously listed under 1.0.0 by mistake) and have since been reworked against a browser-derived conformance corpus - tests/data/layout/capture.html renders each configuration with real CSS and tests/yup_gui/yup_FlexBoxParity.cpp / yup_GridParity.cpp replay the recorded rectangles.
- Breaking:
FlexItem::width,heightandflexBasisnow use-1to mean auto instead of0, matchingGridItemand the min/max fields. A value of0is now a genuine zero size, soFlexItem (component, 0, 0)- which used to mean "auto on both axes" - is a zero-sized item; useFlexItem (component)instead. This is what makesflex: 1 1 0(equal shares regardless of content, the most common flex idiom there is) expressible at all, via the newwithFlexBasis (0) FlexBox: addedwithFlexShrink()andwithFlexBasis()builders; those two fields previously had no fluent setterFlexBox:align-items: stretch- the default - now actually stretches in a single-line (noWrap) container. The line's cross size was only expanded for multi-line containers, so an item with no explicit cross size laid out 0px tall, which is why nested containers with unsized children collapsed. Per CSS,align-contentdoes not apply to a single-line container at all and its one line spans the container's whole cross size; it does still apply to a wrapping container that happens to produce one lineFlexBox:align-items: stretchno longer overwrites an explicit cross size -height: 50pxin a stretch row stays 50px and aligns to the cross startFlexBox:justify-contentno longer double-counts space already consumed byflex-grow. The free space was computed before flexible lengths were resolved and then reused as theflexEnd/centeroffset, pushing items outside the container; it is now recomputed after the flex pass, so a line whose items grow to fill it leavesjustify-contentnothing to shift byFlexBox: flexible lengths are resolved with the CSS §9.7 freeze-and-loop instead of a single proportional pass. An item that hits itsmaxWidthwhile growing (or itsminWidthwhile shrinking) now freezes and the space it could not take is redistributed over the remaining items, so a line no longer leaves a residual gap or overflowFlexBox:align-content: spaceBetweennow accounts for the preceding lines' cross sizes. Middle lines were positioned from the container's start with no accumulated offset, so with three or more lines they overlapped the first one. Everyalign-contentmode now goes through one leading-offset/spacing computationFlexBox: line breaking counts the gaps already consumed on the line. The accumulator tracked item sizes only while the free-space math counted gaps, so wrapped lines overflowed by roughly(n-2) × gapand broke one item too late. Breaking also now uses the hypothetical main size (the base size clamped by min/max) as CSS requiresFlexBox:spaceAroundis no longerspace-evenlyon either axis. CSS puts half a share at each edge and a full share between (extra/2nandextra/n); bothjustifyContentandalignContentusedextra/(n+1)everywhere.spaceBetweenandspaceAroundnow also fall back to flex-start and center respectively when the content overflows, per the box alignment specFlexBox: reverse directions mirror the item's position within the line rather than the final rectangle, somarginLeftstays a left margin inrowReverse(andmarginTopa top margin incolumnReverse) instead of silently acting as its opposite.wrapReverselikewise flips the cross axis rather than only reversing the line order, soalign-itemsresolves against the flipped axisFlexBox:orderis applied with a stable sort, so items sharing anordervalue keep their source order as CSS requiresFlexBox: baseline alignment no longer subtracts the leading cross margin twice, derives the default baseline from the item's clamped cross size, and grows the line when a baseline-aligned item needs more room than the tallest item alone.AlignItems::baselinein a column container now falls back to flex-start, matching browsers - the cross axis there is the inline axis, where a box with no text has no baseline to shareFlexBox: negativeflexGrow,flexShrinkandgapare asserted and clamped to 0 instead of being used as-isFlexBox: removedcalculateLayout(), which was declared but never defined or called, and theLineInfostruct it exposed, whose size fields were copied before being initialisedGrid: fractional (fr) tracks now subtract the gaps before dividing up the remaining space. Track positions advance bysize + gap, so every gappedfrgrid previously overflowed its container by exactly the total gapGrid: placement follows the CSS order - all explicitly positioned items are recorded first, then items locked to a row pick a column within it, then the rest flows. An auto-placed item can no longer claim a cell that an explicitly placed item further down the list owns, and an item that pins only one axis keeps it (a single "either is unset" test previously discarded both)Grid: item sizes and cell sizes are clamped at 0, so margins larger than the cell no longer produce an inverted rectangleGrid: addedGridItem::autoPlaceas the named constant for the automatic-placement sentinel, and documented that YUP's grid lines are 0-based whereas CSS numbers them from 1Grid:calculateTrackSizesmoved out of the class (it never usedthis), and the auto-placement scan is bounded so a pathological span or position cannot allocate without limitFlexBox/Grid:performLayoutno longer rounds to whole pixels.Component::setBoundstakes floats, and rounding position and size independently made adjacent items disagree about the edge they share- Documented that
Grid::TrackInfo::auto_()is a fixedautoRows/autoColumnssize rather than CSS's content-drivenauto, and thatautoRows/autoColumnsalso sizeautotracks inside a template, not just implicit ones. Content-driven sizing needs a measurement hook onComponentthat does not exist yet; the one place the layout engines ask an item for its content size is now a single documented function inyup_FlexBox.cppso it can be swapped without re-plumbing the algorithms
FlexBox feature parity
- Added
spaceEvenlytoFlexBox::JustifyContentandAlignContent, which is an equal share between items and at both edges - distinct fromspaceAround's half-share edges - Added independent
rowGapandcolumnGap.gapis now a shorthand used only for whichever of the two is left at -1; in a row container the columns separate items and the rows separate lines, and in a column container it is the other way round - Added container
paddingLeft/paddingRight/paddingTop/paddingBottomwithsetPadding()shorthands, which shrink the content box every item is laid out in. Padding larger than the target area clamps instead of inverting it - Added CSS
margin: autoviaFlexItem::marginLeftAutoand friends (withAutoMargins()). An auto margin absorbs its share of the free space on its axis beforejustifyContentis consulted, somarginLeftAutopushes an item to the far end of a toolbar and a pair of them centers it. On the cross axis an auto margin positions the item within its line and suppresses stretching - Added
FlexItem::flexBasisPercent(withFlexBasisPercent()), resolved against the container's main-axis size and taking priority overflexBasis
Grid feature parity
- Breaking:
Grid::TrackInfois now a min/max pair (minimum/maximum, each aSizingFunction) instead of the three parallelpixelSize/fraction/isAutofields. All construction still goes through the factory functions, which are unchanged, so only code that read those fields directly is affected - Added
Grid::TrackInfo::minmax(),percent()andfitContent().minmax (px (100), fr (1))is the combinationfralone cannot express: take a share of the leftover space but never drop below 100px. Percentages resolve against the container's full size, not against what is left after the gaps - Track sizing now follows CSS §12.4-12.7: tracks start at their minimum, leftover space grows them towards their maximum, and the fractional tracks then divide up what remains with the same freeze-and-loop as flex-grow - a track whose share would land below its own minimum freezes there and the rest is redistributed
- Added
Grid::repeat (count, track)andGrid::repeatToFill (track, size, gap, defaultSize), the latter being CSS'srepeat(auto-fill, ...). Note thatauto-fitis deliberately not offered: it differs fromauto-fillonly by collapsing tracks that end up with no items, which needs to know each track's contents, so offering both would promise a difference that is not there - Added
Grid::autoFlowwithrow,column,rowDenseandcolumnDense. Placement was previously row-only and sparse-only; a dense flow restarts the cursor for each item so it can backfill the holes a larger item left behind - Added
Grid::justifyContentandalignContent, so a fixed-track grid can finally be centered (or spaced) inside a larger box instead of always leaving the slack at the bottom right - Added
Grid::setTemplateAreas(), mirroring CSSgrid-template-areas, withGridItem::withArea(). The call returns aResultand rejects ragged rows and non-rectangular areas rather than laying out something surprising; a.cell belongs to no area and stays available to auto-placement - Added named grid lines:
Grid::setColumnLineName()/setRowLineName()withGridItem::withColumnStart()/withRowStart(). Names resolve to YUP's 0-based indices at the point they are declared, so the 1-based numbering CSS uses never reaches the placement code - Added
baselinetoGrid::AlignItemsandGridItem::AlignSelf, matchingFlexBox. Items sharing a row align on a synthesized baseline - the item's bottom edge, which is also what a browser uses for a box with no text - Added a
gapshorthand toGrid, sorowGap/columnGapnow default to -1 meaning "use the shorthand", matching the newFlexBoxfields - New
docs/ui/component-layout.mdconcept guide covering both containers, the-1-means-auto convention, and a table of every deliberate divergence from CSS.docs/ui/component-basics.mdno longer claims YUP has no layout manager LayoutDistribution(newlayout/yup_LayoutDistribution.h) is now the single definition of what each alignment mode means, shared byFlexBox'sjustifyContent/alignContentandGrid's. Four separate copies of that switch are what letspaceAroundbe implemented asspace-evenlyin some of them and correctly in none
Python bindings
FlexBox,FlexItem,Grid,GridItem, their enums,Grid::TrackInfo(factory-only, as in C++) andGrid::repeat()/repeatToFill()are now exposed to Python, with the layout containers'itemsandtemplateColumns/templateRowsbound as live arrays. Newpython/tests/test_yup_gui/covers both containers against the same expectations the C++ tests assert. Note that an item only stores a raw component pointer, so a component must be kept alive by the caller for as long as the item referencing it is used
- New GLSL→WGSL direct transpiler in
yup_shading: parses preprocessed GLSL 4.50, lowers GLSL constructs to WGSL equivalents, and emits WGSL 1.0 source. Supports vertex/fragment/compute stages with full builtin mapping, combined sampler splitting, entry-point IO wrapping, and binding assignment matching glslang's SPIR-V assignment 1:1. Does not require SPIR-V or spirv_cross for code generation. Integrated intoShaderTranspiler,ShaderCache, andShaderBundleCompiler. WGSL variants are supported in YSLB bundles via theshader_bundlertool andyup_add_shader_bundle()CMake helper. - New
GpuPipelineclass (rhi/yup_GpuPipeline.h): an immutable compiled render pipeline (vertex + fragment shaders plus fixed pipeline state).compile(ctx, vs, fs, GpuPipelineOptions),compileFromBundle(ctx, ShaderBundle, GpuPipelineOptions), and (whenYUP_ENABLE_SHADER_TRANSPILER = 1)compileFromGlsl(ctx, vertexGlsl, fragmentGlsl, GpuPipelineOptions)all returnResultValue<GpuPipeline::Ptr>. Pipelines carry all the backend-agnostic mirror enums/structs (GpuVertexFormat,GpuPipelineOptions,GpuColorTarget,GpuDepthStencilState, …). - New
GpuFrameclass (rhi/yup_GpuFrame.h): move-only RAII GPU frame scope (GpuFrame::begin(ctx)→submit()→waitForGPU()). Owns the transient GPU resource pools (uniform buffers, texture views, samplers) created while encoding its passes. - New
GpuRenderPassclass (rhi/yup_GpuRenderPass.h): move-only transient render-pass encoder targeting aGpuCanvas. Holds the mutable binding state (setPipeline,setTexture,setUniformBuffer,setVertexBuffer,setIndexBuffer) and encodes draws (draw,drawIndexed,finish). - New
GpuPipelineCacheclass (rhi/yup_GpuPipelineCache.h): thread-safe compile-or-fetch cache forGpuPipelinekeyed by a deterministic SHA1 of the selected native shader sources, entry points, pipeline options, and graphics API. LRU eviction with a configurable entry limit, mirroringShaderCache. - New
GpuBufferclass (rhi/yup_GpuBuffer.h): reference-counted GPU buffer handle wrapping a backend-native GPU buffer.GpuBuffer::create(ctx, GpuBufferType, data, byteSize)uploads immutable vertex/index/uniform data for use withGpuRenderPass. Image::fromTexture(GpuTexture::Ptr): creates anImagewrapping an existing GPU texture (no CPU round-trip). Suitable forGraphics::drawImage().Graphics::drawTexture(GpuTexture::Ptr, Rectangle<float>): draws a GPU texture directly without materialising anImage, avoiding CPU-side ImagePixelData allocation.GpuRenderPassno longer creates a sampler and a uniform buffer per draw. The linear/clamp-to-edge samplers that fill a layout's sampler bindings are created once when theGpuPipelineis compiled, and uniform buffers come from a size-bucketed pool on theGpuDevicethat recycles them when a frame reports GPU completion - so a steady-state workload stops allocating GPU objects after its first frames.GpuFramestays stack RAII; nothing changes for callers.- Fixed the GLSL→WGSL transpiler rejecting comma-separated members in a struct or interface block (
uniform Params { float s, r, rx, ry; }), which failed withExpected ';'. Each declarator now becomes its own member and binds its own array specifiers.
- New glslang (
thirdparty/glslang), SPIRV-Cross (thirdparty/spirv_cross) and SPIRV-Tools (thirdparty/spirv_tools) for shader reflection and cross-compilation (GLSL, ESSL, HLSL, MSL). - New
yup_shadingmodule for cross platform shader handling. - New
ShaderBundleclass (shading/yup_ShaderBundle.h): RIFF binary format (.ysl) that stores original source, per-stage SPIR-V, all transpiled variants (GLSL/ESSL/HLSL/MSL), and fullShaderReflectiondata. Persists to / loads fromOutputStream,File, andMemoryBlockviasaveToStream/loadFromStreamand friends. Lookup by stage + language viafindShader(). - New
ShaderBundleCompilerclass (shading/yup_ShaderBundleCompiler.h): drivesShaderTranspilerto compile + transpile multiple stage/language combinations in one call and returns a fully-populatedShaderBundle. Accepts aShaderBundleCompileRequestwith per-stageShaderBundleEntryitems (stage, target languages,TranspileOptions). - New
BinaryOutputArchive/BinaryInputArchivepair (yup_core/serialisation/yup_BinaryArchive.h): binary stream archives that plug into theSerialisationTraitssystem; used internally byShaderBundleto serialiseShaderReflectiondata intoREFLRIFF chunks. - New standalone
yup_shader_bundlerconsole tool (cmake/tools/shader_bundler): takes a.vertand.fragGLSL (v450 Vulkan dialect) pair on disk and produces a single.yslbundle containing transpiled variants for all target languages (GLSL/ESSL/HLSL/MSL). - New
yup_add_shader_bundle()CMake helper (cmake/yup_shader_bundler.cmake): builds theyup_shader_bundlertool for the host once (cached in the global propertyYUP_SHADER_BUNDLER_EXECUTABLE), runs it at configure time to generate the.ysl, and embeds it into a linkable object library viayup_add_embedded_binary_resources. Works even when the outer build is cross-compiling, since the tool is built in its own host binary tree without forwarding the cross toolchain. Accepts anOPTIONSargument that forwards arbitrary extra flags verbatim toyup_shader_bundler(e.g.--spirv-opt,--target-langs,-DNAME=VALUE,-I<dir>).
- New
yup_dsp_jitmodule (modules/yup_dsp_jit): YDSP, a realtime JIT-compiled audio DSP language compiled to native machine code viaasmjit_library(x86-64 and AArch64). Full compiler pipeline (lexer, parser, type system with realtime-safety enforcement, optimiser, AsmJit backend) plus a zero-allocation realtime runtime (YdspCompiler,YdspAudioGraph). Supports Faust-style composition algebra and Cmajor-styleprocessor/graphdefinitions, sample and block processing, history state, parameters and meters, and sidechain/scratch buffers. YdspCompiler::compile()gains an optionalimportBasePathargument:importdirectives in a patch now resolve relative to that directory (the patch's folder) instead of the process working directory, and nested imports inside an imported file resolve against that file's own folder. This makes multi-file patches loadable from disk; the "YDSP Synths" demo passes each patch's path and ships five importable effect processors indata/synths/fx/(Delay,Compressor,Reverb,Distortion,Chorus), one wired into the graph of each demo synth.- Closed a set of silent-failure gaps found by an audit against the language spec:
min/max/clamp/abs/signgain dedicated integer opcodes (minI/maxI/clampI/absI/signI, branchless on both asmjit backends, compare+select on wasm) instead of only accepting float operands; endpoint annotations ([[ key: value ]]) now go through a whitelist that warns on an unrecognized key, matching every other annotation scope, andunit/step/styleare plumbed all the way toYdspParameterInfoand the YDSP Synths demo's slider setup;stream[N]withN != 1is now a compile error instead of silently yielding mono;buf[i] += xands.field += x(and every other compound-assignment operator, including the previously-missing/=/%=) now desugar correctly via a deep-cloned target instead of failing with "Unknown symbol ''"; the lexer accepts.5,1.,0x1F,0b1010and1_000literal forms and\n/\t/\"/\\string escapes; automating a non-float32parameter is now counted ingetDroppedEventCount()instead of silently discarded; and the function inliner has a re-entrancy guard, so recursion the analyzer's own check misses now fails with a diagnostic instead of overflowing the host process stack. - The one-per-test
YdspCompiler/YdspAudioGraphrecompiles inyup_YdspGraphTests.cpp'sYdspElectricPianoTestsfixture andyup_YdspExamplePatchTests.cpp's two shipped-patch sweeps are replaced with a compile-once cache (yup_YdspTestPatches.h'scachedPatch/restoreFreshState) and a single merged sweep, cutting redundant JIT compiles in the test suite.
- New
yup_aimodule (modules/yup_ai): LLM client and AI integration classes depending onyup_coreandyup_events.
LLMClient(yup_LLMClient.h): abstract base for chat-completion backends withcomplete()andcompleteStreaming()methods, tool-call loop support viarunToolLoop(), and structured output viaLLMSchemaJSON Schema or GBNF grammars.LLMHttpClient(yup_LLMHttpClient.h): HTTP transport forLLMClientwith retry and timeout logic, handling streaming SSE and non-streaming JSON responses.LLMClientFactory(yup_LLMClientFactory.h): creates the correctLLMHttpClientsubclass fromLLMClient::Options::provider, with convenience factories for each provider.LLMMessage(yup_LLMMessage.h): chat message with four roles (system, user, assistant, tool), optional tool calls, and serialisation to/from OpenAI ChatML JSON.LLMResponse(yup_LLMResponse.h): parsed completion response with choices, token usage, tool-call extraction, streaming chunk accumulation, and error handling.LLMTool(yup_LLMTool.h): callable function descriptor with JSON Schema parameters and a local handler, serialised to OpenAI function-calling format.LLMToolRegistry(yup_LLMToolRegistry.h): thread-safe registry forLLMToolinstances with snapshot, lookup, dispatch, and tools-array serialisation.LLMSchema(yup_LLMSchema.h): fluent builder for JSON Schema objects (string,number,integer,boolean,array,object,oneOf) used in structured-output requests across all providers.
LLMOpenAIChatClient(yup_LLMOpenAIChatClient.h): OpenAI Chat Completions API - also compatible with Ollama, DeepSeek, OpenRouter, and llama-server.LLMOpenAIResponsesClient(yup_LLMOpenAIResponsesClient.h): OpenAI Responses API (GPT-5+, reasoning models).LLMAnthropicClient(yup_LLMAnthropicClient.h): Anthropic Messages API (Claude models).LLMGeminiClient(yup_LLMGeminiClient.h): Google Gemini generateContent API.
EmbeddingModel(yup_EmbeddingModel.h): OpenAI-compatible HTTP embedding model withembed()/embedBatch()andcosineSimilarity()helper.
MCPTypes(yup_MCPTypes.h): JSON-RPC 2.0 request/response/error types, MCP capability flags, tool and resource definitions withtoVar/fromVarserialisation.MCPTransport(yup_MCPTransport.h): abstract transport interface for JSON-RPC messages (stdio, HTTP/SSE, sockets, in-process).MCPClient(yup_MCPClient.h): synchronous MCP client withinitialize()handshake,listTools()/callTool(),listResources()/readResource(), and tool-import bridgeregisterToolsWith().MCPServer(yup_MCPServer.h): MCP server exposing local YUP tools and resources over a transport, withregisterTool()/registerResource(),start()/stop(), and placeholderstartStdio()/startHttp().
- Python bindings for
yup_ai(modules/yup_python/bindings/yup_YupAi_bindings.cpp): exposes LLM client, provider, messages, tools, responses, MCP types, client, and server to Python via pybind11.
-
ComponentEffect.displayToContent/contentToDisplayandComponent.getChildPointFromLocal/getLocalPointFromChildare overridable from Python; the point mapping helpers,setManuallyComposited()andrenderToTexture()are bound. -
GpuCanvas.beginDraw()accepts an optionalscale, andComponent.renderToTexture()an optionalscale. -
The new
Componentcallbacks are overridable from Python:hitTest,childBoundsChanged,parentSizeChanged,indexInParentChildrenChanged,focusOfChildComponentChanged,keyStateChangedandmodifierKeysChanged.hitTestis also callable directly, and the newyup.FocusChangeTypeenum is bound alongside the extracauseargument onComponentNative.setFocusedComponent() -
Bound the rest of
Component's public surface:getSafeAreaBounds(),setPaintProfilingDisabled()/isPaintProfilingDisabled(),setMetric()/getMetric()/findMetric(),setCachedToTexture()/isCachedToTexture(),setComponentEffect()/getComponentEffect(),addComponentListener()/removeComponentListener()andsnapshotToImage()/snapshotToTexture(). The effect and listener methods needed somewhere for their argument to come from, soComponentEffect(whoseapply()is the subclass point),ComponentListenerandComponentPaintMetricsare bound too, andaddComponentListener()keeps the Python listener alive - the C++ listener list only holds weak references.PyComponentalso gained theopacityChangedoverride it was missing, the oneComponentvirtual no Python subclass could override -
The drag-and-drop payload is bound as
yup.DragAndDropData(withFiles()/withText()/withUris()builders,getFiles()/getText()/getUris()getters andhasFiles()/hasText()/hasUris()/isEmpty()predicates), and the fiveComponenthooks that carry it -isInterestedInDrag(),itemsDropped(),itemDragEnter(),itemDragMove()anditemDragExit()- are callable from Python. A PythonComponentsubclass could already override those hooks, but no drop could ever reach them: the payload had no Python type, so handing it back to the override threw the moment the platform delivered a drag.withFiles()takes any iterable ofFile, because theArray<File>binding's name is derived from the typeid at runtime and is not something to require of callers -
yup.GpuCanvas.create()now accepts the three-argument formcreate(ctx, width, height). pybind11 does not see C++ default arguments, so the bound signature demanded the optionalclearColortoo and every call raisedTypeError: create(): incompatible function arguments. -
yup.GpuCanvas.beginDraw()returns a borrowing wrapper of the canvas'sGraphicsinstead of trying to copy it.Graphicsis non-copyable, so pybind11's defaultcopyreturn policy made every call raiseRuntimeError: return_value_policy = copy, but type yup::Graphics is non-copyable!; the wrapper now borrows the object and keeps the canvas alive alongside it. -
Widget style identifiers are now reachable from Python:
Label's nested C++Stylestruct is bound asyup.Label.Style, solabel.setColor (yup.Label.Style.backgroundColorId, yup.Colors.darkblue)works instead of raisingAttributeError: type object 'yup.Label' has no attribute 'backgroundColorId'. The struct holds only static members and is deliberately not constructible -
yup.Colornow converts implicitly toyup.GpuColor, soyup.Colors.blackcan be passed anywhere aGpuColoris expected (GpuRenderOptions(True, yup.Colors.black)) instead of raising aTypeError. C++ already converts implicitly -GpuColor's constructor accepts any type with float-component accessors - so the binding registers that constructor and the conversion that uses it -
Justification::Flagsmembers are now implicitly convertible toJustification, soyup.Justification.leftcan be passed straight to any API taking aJustification(Graphics.fillFittedTextand friends) instead of raising aTypeError. Combining flags with|yields the underlying integer, so anintconverts too -
An exception raised on a background thread no longer stops the dispatch loop.
paint()runs on the render thread, andPyErr_CheckSignals()does nothing off the main thread, so the message-thread signal timer is the only thing that can ever notice Ctrl+C - stopping the loop from the render thread took that timer down with it and left the process deaf to SIGINT for the rest of its life. Only the message thread stops the loop now; a failingpaint()reports and keeps running, so Ctrl+C and the window's close button still work. -
Bound
ApplicationTheme:yup.ApplicationTheme.getGlobalTheme()plusgetDefaultFont(),getDefaultIconFont()andgetDefaultMonospaceFont(), so Python code can obtain theme fonts the same way C++ does.getGlobalTheme()raises when no theme is set instead of returning a null pointer -
Exceptions raised in Python overrides that C++ calls back into (
paint,timerCallback,messageCallback, ...) now reach the application'sunhandledException()with the original type and traceback, rather than terminating the process.START_YUP_APPLICATION'scatchExceptionsAndContinuenow governs the no-override fallback: when true the dispatch loop absorbs the error and keeps running, when false it stops so control returns to the interpreter. Previously that path calledstd::terminate()unconditionally. -
TestApplication(theyup.TestApplicationcontext manager used by the Python test suite) now keeps a reference to the application it constructs. The onlypy::objectwas a local in the scope's constructor, so the application was destroyed as soon as construction finished andYUPApplicationBase::getInstance()was null for the entire test. Everything guarded by it silently did nothing — most visiblysendUnhandledException(), so exceptions raised in Python overrides never reachedunhandledException(). The scope now also callsshutdownApp()on teardown, which it never did. -
Python exception reporting no longer prints each error's location twice. Both
Helpers::printPythonExceptionandPyYUPApplication::unhandledExceptioncombinederror_already_set::what()withtraceback.print_tb(), butwhat()already renders the traceback itself, so every error appeared once in pybind'sAt: file(line): funcform and again in Python's. Both now usetraceback.print_exception(), which produces the single rendering the interpreter would. -
Ctrl+C is now noticed promptly instead of only when the application next happens to run Python on the main thread (which for a GUI whose only Python code is
paint()on the render thread meant the next window focus or mouse event). The dispatch loop runs with the GIL released and CPython only runs signal handlers on the main thread, sorunApplicationnow starts a message-thread timer that callsPyErr_CheckSignals()everymessageManagerGranularityMilliseconds— a parameter that was passed but never used. It is a timer rather than a slicedrunDispatchLoopUntil()because on Apple platforms that overload does not invoke the event loop callback that pumps SDL. -
Ctrl+C stops a Python application again. Now that the dispatch loops catch rather than letting the
KeyboardInterruptunwind out ofrunDispatchLoop(),unhandledExceptionhas to stop the loop itself — it previously only setcaughtKeyboardInterruptand returned, leaving the loop running.runApplicationalso no longer re-raises a signal it has already reported. -
Exception reporting no longer raises from inside the handler when the traceback is null, which is the case for a
KeyboardInterruptdelivered byPyErr_CheckSignals()at a C boundary — a nullpy::objectcannot be passed to a Python call and threwcast_error. -
PyYUPApplication::unhandledExceptionno longer dereferences a nullstd::exception*.sendUnhandledException()passes null for acatch (...), which the newly addedYUP_CATCH_EXCEPTIONsites make reachable. -
Destroying a Python
Componentthat is on the desktop no longer deadlocks the process. Tearing the native peer down joins the render thread, which can be blocked acquiring the GIL insidepaint()/refreshDisplay()- and Python holds the GIL while it destroys the component, so neither thread could proceed and the message thread wedged inside the SDL event pump, taking the window's close button, the dock icon and Ctrl+C with it. ThePyComponenttrampoline destructor now detaches from the desktop with the GIL released, andaddToDesktop()/removeFromDesktop()are bound with agil_scoped_releasecall guard. -
unhandledExceptionreporting no longer touches Python from the thread that threw. Reporting from the render thread meant acquiring the GIL with no abort path while the message thread could be holding it to join that same thread, and after an exception it ran once per frame -import traceback, format, print - which is what made the deadlock above reproducible. Off the message thread the reporting and the stop request are now marshalled to it as a single message, so a stop can never latch the loop shut before the report is dispatched, and the throwing thread touches no Python at all.reportUnhandledExceptionis alsonoexceptthroughout: it is always called from aYUP_CATCH_EXCEPTIONhandler running in a C or Objective-C frame, where a raisingunhandledExceptionoverride would previously escape and leave the platform event locks permanently held. -
An exception on the render thread now stops the application instead of repeating forever.
unhandledExceptionreturned early on any thread that was not the message thread, so withcatchExceptionsAndContinue=Falsea failingpaint()was reported on every frame and the app never quit;stopDispatchLoop()is safe to call from any thread and is now used. -
Ctrl+C no longer throws across the platform timer's C frame. The signal-check timer captured the raised error and stops the loop in place;
START_YUP_APPLICATIONre-raises it as a normalKeyboardInterruptafter the application has shut down.START_YUP_APPLICATIONalso always shuts the application down now - previously a re-raised exception or a signal checked after the loop skippedshutdownApp()entirely. -
PyYUPApplication::unhandledExceptionimported__builtins__, which is not a module insys.modules, so the branch handling a non-Python exception raised from inside the exception handler. It now usesPYBIND11_BUILTINS_MODULE. -
The
yup_rhibindings now cover the surface the RHI grew: textures (GpuTexture.create/upload,GpuTextureDesc,GpuTextureViewDesc,GpuTextureDataDesc), samplers (GpuSampler,GpuSamplerDesc,GpuRenderPass.setSampler), the full render-pass API (setColorAttachment,setDepthStencilAttachment,setResolveTarget,setViewport,setScissorRect,setStencilReference,setBlendColor,GpuDepthStencilOptions), vertex layouts and multi-target pipeline options, compute (GpuComputePipeline,GpuComputePass,GpuWorkgroupSize), theGpuTargettexture and view overloads withreadPixels(), and theGpuDevicecapability probes and buffer create/update/read entry points -
Fixed
GpuPipeline.compile*being uncallable from Python: they returnResultValue<T>, which has no registered Python type, so every call raised at return-conversion time. They now unwrap and raise the compiler diagnostics as a Python exception.GpuTarget.beginRenderPasswas not bound at all, andGraphicsContexthad no class binding, soGraphics.getGraphicsContext()raised as well - between them, neitherpython/demos/gpu_triangle.pynorgpu_effects.pycould run -
GpuRenderPass.setUniformBufferand the buffer entry points take any buffer-protocol object (bytes,bytearray,memoryview, numpy arrays) instead of onlybytes, and size them in bytes rather than items -
GpuTextureFormatexposes all 25 formats (5 were bound), andGpuVertexFormatall 17 (7 were bound). AddedGpuColorWriteMask(composable with|and&),GpuFilter,GpuWrapMode,GpuTextureType,GpuTextureAspect,GpuTextureViewDimension,GpuBufferType.storageand the remaining blend factors -
GpuFrameandGpuRenderPasskeep the objects they borrow alive (py::keep_alive), so a pass can no longer outlive the frame whose pools it points into -
GpuShaderSourceis now exposed, now that its blob fields own their data:code,bindingMapandglFixupare bytes-in, bytes-out properties accepting any buffer-protocol object.GpuPipeline.compileandGpuComputePipeline.compileare bound alongside the existingcompileFromGlsl, so Python can compile from native shader sources without going through the GLSL transpiler.GpuCanvas.beginDraw()gained an optionalGpuFrameDescriptorargument (also newly bound, along withGpuDitherMode), giving Python control over msaa/dither/loadOp/clearColor for the offscreen 2D frame it opens - bound as two overloads rather than one defaulted argument, sinceregisterYupGraphicsBindingsruns beforeregisterYupRhiBindingsand a default value referencingGpuFrameDescriptorat bind time would throw at import -
Fixed
registerYupRhiBindingsbeing called underYUP_MODULE_AVAILABLE_yup_graphicswhile its translation unit compiles underYUP_MODULE_AVAILABLE_yup_rhi, which silently dropped the RHI bindings in a graphics-less build -
Added
python/demos/gpu_cube.py: a textured, depth-tested spinning cube driving vertex and index buffers, an uploaded texture and a sampler entirely from Python -
python/demos/gpu_cube.pyno longer renders its cube inside out. Its face corners wound counter-clockwise seen from outside, but the RHI bakes a clip-space Y-flip into the vertex stage — the GL backend invertsglFrontFaceand Vulkan relies on naga'sADJUST_COORDINATE_SPACEto compensate for it — soGpuCullMode.backculled the side facing the camera instead of the far side. The corner order now matchesSpinningCubeDemo'skCubeVerts -
Fixed windows never appearing when a YUP application is run from a Python interpreter on macOS. A bare interpreter is not a bundled app, so the process starts as a non-UI one; SDL would normally set the activation policy, but it only does so when it is the one to create
NSApp, and YUP'sMessageManagergets there first.START_YUP_APPLICATIONnow transforms the process to a foreground application up front. Deliberately not done ininitialiseYup_Windowing(), so a plugin hosted in a DAW can never transform its host's process -
Added
ColorGradientbindings (plus itsType/Spreadenums and nestedColorStop). The previous binding had been commented out when the C++ API replacedbool isRadialwith aTypeenum, which leftGraphics.setFillColorGradient()/setStrokeColorGradient()bound but uncallable -
Component::paintSubtreenow clears itsisRepaintingflag through a scope guard. Apaint()override that throws - which is how a Python error surfaces - previously skipped the reset, leaving the flag set permanently so every laterrepaint()of that component tripped an assertion pointing at the wrong cause -
Bound the
yup_corefacilities that had no Python surface at all:Logger(withsetCurrentLogger()accepting a Python logger object),FileLogger,DynamicLibrary,SHA1,CancelToken/CancelTokenSource,WaitableTimer,StringPool,TextDiff,DynamicObject,AbstractFifo/SingleThreadedAbstractFifo,Expression(plus itsExpressionScope),LocalisedStrings,YAML(plusFormatOptions/Spacing),IPAddress,MACAddress,NamedPipe,WebInputStreamand theInputSourcefamily -
registerStatisticsAccumulator()lets a C++ statistics type be addressed from Python the way the other generic templates are, which is what makesyup.StatisticsAccumulator[float]resolve instead of raising. Onlyfloatis registered: Python has a single float type, so registeringdoubleas well would map both spellings onto the same key and overwrite the first rather than adding an overload -
Bound
Fitting,CubicBezierandDrawableinyup_graphics;KeyModifiers,KeyPress,MouseWheelData,ProgressBarandSwitchButtoninyup_gui; andMessageBase,Message,CallbackMessageandMessageListenerinyup_events. The two widgets needed trampolines for the same reason the rest do -paint()is the subclass point -
Fixed
ImageFormatManager::createReaderForleaking the stream it opened. The manager opens the file and hands the reader anInputStreamto own, butcreateReaderForreleased the Python wrapper of that stream on the way in while the format it built had no way to adopt it, so every call leaked oneFileInputStream.ImageFormatReadernow has a constructor that takes the caller's stream withdeleteSourceWhenDestroyedset, and the bindings hand a Python format the stream the manager opened, so the reader deletes it exactly as the C++ path does -
Methods that take a
std::unique_ptr<T>parameter are only callable with a Python-constructedTwhenTis registered withpy::smart_holder: pybind11 3.x moves the object out of the wrapper and disowns it, which it cannot do for aunique_ptrholder.InputSource(withFileInputSourceandURLInputSource),XmlElementand the wholeInputStreamhierarchy (FileInputStream,MemoryInputStream,BufferedInputStream,SubregionStream,GZIPDecompressorInputStream,WebInputStream) were migrated, soXmlDocument.setInputSource(),XmlElement.addChildElement()andZipFile.Builder.addEntry()now transfer ownership instead of releasing a wrapper nothing accounted for. The trampolines for those hierarchies also carrytrampoline_self_life_support, which is asmart_holderrequirement - pybind11 rejects the combination at compile time otherwise -
MessageBase,MessageandCallbackMessageare registered on their naturalReferenceCountedObjectPtrholder instead of the defaultunique_ptr.MessageBase::post()andMessageListener::postMessage()manage the message themselves - the queue takes its own reference, and the failure path only deletes a message whose count is still zero - so the previous bindings had to hand ownership over throughrelease()to stop Python deleting a message the queue still held. With the refcounted holder both methods bind directly, and a message Python still holds stays valid after posting -
These methods now consume the Python object they are given: after
addChildElement(),setInputSource(),addEntry()orZipFile(stream), using the wrapper raisesValueError: ... Python instance was disowned. That is the ownership the C++ API documents - the container deletes what it was given - and it replaces the previous behaviour, which leaked the reference instead -
The new surface is covered by additions to
python/tests/, anddocs/scripting/python-bindings-coverage.mdrecords the per-module inventory of bound against declared API that these gaps were found from -
Bound
MouseEvent,MouseListenerandTextInputTarget, and declared theMouseListenerbase ofComponent. No PythonComponentsubclass could receive a mouse callback that carried an event:mouseMove(),mouseDrag(),mouseUp(),mouseDoubleClick()andmouseWheel()were already routed to overrides, butyup.MouseEventdid not exist, so the moment the platform delivered one, the conversion back to Python threw.Component::addMouseListener()now also keeps the Python listener alive, matchingaddComponentListener(), and the wheel trampoline looks its override up under the C++ namemouseWheel- it asked formouseWheelMove, a name nothing in the tree defined, so a wheel override written the obvious way was silently never called -
Bound the
yup_guiinput and widget types the coverage page still listed as missing:ScrollBar,ListBox,ListBoxModel,ListBoxItem,ComboBoxandTextEditor, each with its nested enums (ScrollBar::Orientation/VisibilityMode,ListBox::Orientation/SelectionMode,ListBoxItem::IconPosition) and its nestedStylestruct exposed for its theme identifiers the wayyup.Label.Stylealready was.ListBoxModelis the subclass point for a Python list model and is dispatched through a trampoline;setModel()pins the model, which the ListBox never owns.TextEditorneeded its own trampoline forgetTextInputRect(), theTextInputTargetvirtual it implements -
ListBoxItem::setIconDrawable()/getIconDrawable()andListBoxModel::refreshComponentForRow()are deliberately not bound: the first pair traffic instd::shared_ptr<Drawable>whileDrawableuses pybind11's defaultunique_ptrholder, and the second hands theListBoxownership of the component it returns, which cannot be taken away from a Python-owned instance without inviting a double free.setIcon()andpaintListBoxItem()/getRowText()/getRowIcon()cover the same ground; both exclusions are recorded on the coverage page -
Bound
AudioIODeviceTypeandAudioDeviceManager::getAvailableDeviceTypes(), whichpython/demos/audio_device.pyneeds to list what the machine offers — the demo raisedAttributeError: 'yup.AudioDeviceManager' object has no attribute 'getAvailableDeviceTypes'. The manager owns its device types, sogetAvailableDeviceTypes()returns a fresh Python list whose entries borrow from it: a caller keeping a type keeps the manager alive alongside it, andcreateDevice()hands Python the device it creates withtake_ownership, which is what the C++ contract asks for.AudioIODeviceType::Listener,addListener()andremoveListener()stay unbound for want of a trampoline -
PositionableAudioSourcewas bound without declaring itsAudioSourcebase, so pybind11 treated the two as unrelated types:python/demos/audio_player.pydied withTypeError: setSource(): incompatible function arguments ... (self: yup.AudioSourcePlayer, newSource: yup.AudioSource) ... Invoked with: ..., <yup.AudioTransportSource object>. An audit of all 289py::class_registrations against the module headers found this to be the only such omission (a scan that has to match the two-phasepy::class_<...> classX (m, "X")form, since the single-phase declarations are all correct).AudioSourcePlayer::setSource()andAudioTransportSource::setSource()now also pin the source they are given withpy::keep_alive: both C++ contracts say the object playing it does not own it, and nothing in Python could otherwise express that -
AudioFormatReaderSourceno longer takesdeleteReaderWhenThisIsDeletedfrom Python, and no longer reads a freed reader.createReaderFor()hands Python astd::unique_ptr, and pybind11 cannot take a raw-pointer argument's ownership away from a live wrapper, soAudioFormatReaderSource(reader, True)— whatpython/demos/audio_player.pyandaudio_player_waveform.pyboth did — built a source that deleted a reader Python was about to delete too, and that read the reader after Python had dropped it. The source now always borrows and is pinned to the reader withpy::keep_alive, so the reader outlives every use of it (C++ callers wanting the transfer still use thestd::unique_ptrconstructor). The symptom this fixes is not a crash: a dead reader reports a total length of 0,AudioTransportSource::hasStreamFinished()compares that with the read position,0 >= 0, andgetNextAudioBlock()clears the transport'splayingflag on the first block, so the transport went silent andisPlaying()never became true -
yup.AudioBufferis subscriptable, soyup.AudioBuffer[float]resolves toyup.AudioBufferFloatthe wayyup.Rectangle[float]andyup.StatisticsAccumulator[float]already did —python/demos/audio_player_waveform.pyfailed withTypeError: type 'yup.AudioBufferFloat' is not subscriptable. The existing callable alias keeps working, so a subscription was added to it rather than replacing it with the type-keyed dictionary the other templated types expose: Python has one floating-point type, so such a dictionary could holdfloatalone and the double specialization stays reachable asAudioBufferDouble -
python/demos/audio_player_waveform.pyno longer glitches while it plays. Itspaint()read the whole file through the sameAudioFormatReaderthe audio thread was pulling, about 800 reads per frame:AudioFormatReader::read()re-seeks its stream and allocates on every call and the reader takes no lock, so the two threads moved the shared stream position out from under each other, and the audio thread competed for the allocator with a storm of render-thread reads. The peak envelope is now computed once, before playback starts, and painting only reads that list -
python/demos/layout_flexgrid.pylays its panels out again. EveryFlexItemwas built with an explicit width of 0, and a literal 0 is a real zero size rather than "auto", soalign-items: stretchskipped all five labels and the window showed nothing but the component's own black background. The sidebars were also added to a nestedbodyFlexthat never hadperformLayout()called on it, andself.contentwas added to both boxes - the nested row now runs over the band the column box leaves for the content, in aRectangle[float]sinceComponent::getLocalBounds()returns floats andRectangleIntrejects them -
python/demos/layout_rectangles.pyruns at all now. Itspaint()had never executed past its first statement:w - 80measured from float component bounds was fed toRectangle[int], the eight.to<float>()conversions are the C++ spelling of the PythontoFloat(),Graphics::drawTextdoes not exist (fillFittedText (text, font, rect, justification)is the API, and the font comes fromApplicationTheme),Justification::centredbecameJustification::center, and the greys aredarkgray/gray/lightgray- JUCE's British spellings were not carried over. It also carves its frame out of the window bounds withremoveFrom*instead of subtracting fromwandh: the subtraction went negative once the window was narrower than the 40px margin, andremoveFrom*asserts on a negative extent (jlimit (0, extent, delta)inyup_Rectangle.h), so resizing the window down to zero width tripped it -
Note for anyone else driving widgets from Python: laying widget text out goes through
ApplicationTheme(ComboBox::updateDisplayText(),TextEditor's styled text andListBoxItem::calculateLayout()all ask the theme for a font), and the global theme only exists while an application is initialised. Outside one those lookups dereference a nullReferenceCountedObjectPtrand take the process down, so a widget that carries text has to be created and used inside a running application - in the suite that is thejuce_appfixturetest_ApplicationTheme.pyalready used, which the new widget tests take as well -
python/demos/matplotlib_integration.pyis an actual port of popsicle's demo now, instead of a chart drawn out of YUP primitives.make_plot()/generate_plot_png()build the linear-regression figure in a child process - matplotlib is not thread safe, so it stays off the UI thread - and hand the PNG back over amultiprocessing.Queue; a 24Hzyup.Timerpolls that queue, decodes the bytes withImage.loadFromData()into a child component that paints them, and fades that child in withsetOpacity()while a star spinner turns behind it. Four YUP-for-JUCE substitutions were needed: the chart widget is a plainComponent(YUP has noDrawableImage), the fade is driven from the timer (noDesktop::getAnimator()),fillAll()takes no color so the white fill issetFillColor()+fillAll(), and the child is added withaddAndMakeVisible()because a YUP component starts out withisVisible() == false, where JUCE's starts visible. Two of those were only found by running it. The image child must not be opaque:Component::hasOpaqueChildCoveringArea()ignores the child's opacity, so an opaque child covering the parent makesinternalPaint()skip the parent'spaint()entirely and the window showed nothing but black - no white background, no spinner - until the chart arrived. And the timer is a plainChartPoller(yup.Timer)holding the component rather than a second base of it, because a Python class deriving from two bound YUP classes (Component+Timer, as the original'sMainContentComponent(juce.Component, juce.Timer)is) crashed with a badthisinsidePyComponent's trampoline destructor when the window closed - the dealloc walk of such an instance runs pybind11's multiple-inheritance value_and_holder bookkeeping, and the class had been registered without the pair the pybind11 documentation asks of every trampoline. The Gui trampolines now derive frompybind11::trampoline_self_life_supportand their classes are registered withpy::smart_holder, which is what the Core and Graphics bindings had already been doing for their own trampolines.python/tests/test_yup_gui/test_MultipleInheritance.pycovers it: the two-base case runs in a child interpreter, because the failure mode is a signal rather than an exception, and the test fails if the child is killed by one instead of reporting an error.Image.loadFromData()raisesValueErroron a payload it cannot decode, where theImageCache::getFromMemory()it replaces returned a nullImage
- The graphics synthesizer example's oscillator displays now draw the waveform the voices play with a
yup_rhifragment shader: a raymarched teal ridge landscape whose far ridge is one period of it, animating slowly while shown. The pipeline is compiled once and shared by both displays, and the vector display remains the fallback without a GPU. - The graphics synthesizer example's oscillator display can be drawn on freehand. The stroke is analyzed into 64 phase-preserving partials, which the PARTIALS view then edits.
Component3DDemopresents an interactive widget panel on a curved 3D surface;WidgetsDemogains draggable corner handles applying anAffineTransformto the widgets panel; the Wave effect inComponentEffectsDemonow maps input, with widgets inside the effected area.Component3DDemorenders the panel texture and its 3D scene at the display scale.- Fixed the Python demo aborting on "Run Python!":
PythonDemois now bound withpy::smart_holder, matching itsComponentbase. SpinningCubeDemoexample (examples/graphics): rewritten to the new RHI shape -GpuFrame+GpuCanvas::beginDraw+GpuRenderPassfor both the indexed cube draw and the separable two-pass blur (H+V sharing oneGpuFrame),isGpuAvailable()capability probe, and live GLSL editing viaGpuPipeline::compileFromGlsl. The default Lottie animation is now played back per-frame into an offscreenGpuCanvas(2D path) and sampled by the cube's fragment shader so the animation is texture-mapped onto every cube face.WidgetsDemoexample (examples/graphics/source/examples/Widgets.h): the placeholder image button is now a workingImageHitTestButtondemonstratingComponent::hitTest- it drawsdata/logo.pngand samples the image's alpha at the hit point, so only the logo's opaque pixels are clickable and the transparent ones fall through to what is behind. The hover highlight goes through the same test, so moving the pointer over a transparent region inside the button's bounds drops it.SpinningCubeDemoexample (examples/graphics): rewritten to the new RHI shape —GpuFrame+GpuCanvas::beginDraw+GpuRenderPassfor both the indexed cube draw and the separable two-pass blur (H+V sharing oneGpuFrame),isGpuAvailable()capability probe, and live GLSL editing viaGpuPipeline::compileFromGlsl. The default Lottie animation is now played back per-frame into an offscreenGpuCanvas(2D path) and sampled by the cube's fragment shader so the animation is texture-mapped onto every cube face.AIDemoexample (examples/graphics/source/examples/AI.h): interactive demo for all four LLM providers (OpenAI Chat, OpenAI Responses, Anthropic, Gemini) with model and API key configuration, system prompt editing, streaming and non-streaming completion, tool calling, MCP server integration, and embedded text generation.ArtboardDemoexample (examples/graphics/source/examples/Artboard.h): the loaded Rive artboard now queries a named node viaArtboard::findNode(name, type, bounds shown in a status label) and attaches a rectangular marker component to it withArtboard::attachComponentToNode— a "Marker" combo switches between filling the node's bounds and tracking only its position, "Pivot" and "Anchor" combos choose which component point lands on which node point in track mode, and an "Apply transform" toggle rotates the marker with the node — and the marker follows the node on reflows and resizes.- New
ArtboardLayoutDemoexample (examples/graphics/source/examples/Artboard.h, registered as "Artboard Layout"): sharesArtboardDemoBase's controls withArtboardDemobut loadsdata/layout-ui.rivand attaches a live, nestedArtboard(playing the file'sKeyboardartboard) to thekeyboard_slotlayout placeholder in the mainWireframeartboard, instead of a plain marker rectangle. ArtboardDemoexample: the displayed Rive file can now be replaced at runtime by dropping a.rivfile onto the demo, which rebuilds the artboards from the dropped file while keeping the current fit, alignment and marker settings. The demo outlines itself while a.rivis dragged over it.AudioExampleexample (examples/graphics/source/examples/Audio.h): reworked into a Vital-style instrument. Each oscillator gets a display that either draws the reconstructed waveform or edits its partials as draggable magnitude bars, writing sine coefficients the wavetable, sync and morphing backends all render; the drag publishes at most one generation bump per UI frame so it cannot queue an inverse FFT per mouse event across every sounding voice, and the reconstruction's peak is measured on the message thread and applied as a coefficient scale so an edited spectrum cannot exceed full scale. Unison adds up to five detuned, stereo-spread slots per oscillator, built from bareWavetableOscillatorsatellites rather than furtherSynthOscillatorcopies (which own four wavetable oscillators each once sync and morphing are counted) and offered on the wavetable algorithm alone, where they are exact. A DAHDSR envelope with a drawn curve replaces the singleSmoothedValuefade, and the voice now renders stereo. The demo also plays the first available hardware MIDI input, collected through aMidiMessageCollectorinto the same bufferMidiKeyboardStatereads, so external notes light up the drawn keys as well as sounding. Fixes theDetunecontrol, which the UI wrote but the engine never read, so both oscillators always played the same frequency.
- Added
yup_add_bundled_resources()and aBUNDLE_RESOURCESargument onyup_standalone_app(), taking<file>@<relative-dest>pairs (same convention asPRELOAD_FILES) and placing them whereFile::getSpecialLocation (File::bundleDirectory)can find them at runtime:Resources/<relative-dest>in the app/plugin bundle on Apple,app/src/main/assets/<relative-dest>on Android, and preloaded into the Emscripten virtual filesystem examples/graphics: the Rive artboard and Lottie demos now bundle their.riv/.lottiefiles viaBUNDLE_RESOURCES(Android, iOS, Emscripten) and read them back throughFile::getSpecialLocation (File::bundleDirectory), instead of compiling them into the binary withyup_add_embedded_binary_resources()- justfile recipes now use per-platform build directories (
build/mac,build/ios,build/android,build/emscripten,build/ninja,build/win), so switching platforms no longer requiresjust cleanand preserves downloaded FetchContent dependencies per platform. Thejust buildrecipe gains aPLATFORMparameter (defaultmac). yup_standalone_appgains aMAXIMUM_MEMORYEmscripten argument (-sMAXIMUM_MEMORY); when set it caps the heap thatALLOW_MEMORY_GROWTHmay reach.yup_testswasm build: raisedINITIAL_MEMORYto 256 MB, addedMAXIMUM_MEMORYcap of 1 GB, and reducedSTACK_SIZEto 1 MB to give the heap room for concurrent pthread stress tests; fixesRuntimeError: memory access out of boundsin CI.- Fetched third-party dependencies (SDL3, Perfetto, plugin SDKs) are now cloned shallowly (
--depth 1) and skip network update checks on reconfigure, speeding up fresh configures and reconfigures. Shallow cloning is automatically disabled when aGIT_TAGis a commit hash. - Android: full 16 KB page size compatibility — generated Gradle projects bumped to AGP 8.5.2 / Gradle 8.7 (uncompressed native libraries are zip-aligned to 16 KB),
jniLibspackaging made explicitly non-legacy, andndkVersionpinned to r27c (overridable viaNDK_VERSION), which ships a 16 KB-alignedlibc++_shared.so. CI NDK updated to r27c accordingly. Application shared libraries were already linked with-Wl,-z,max-page-size=16384. - The vendored Rive runtime now builds with its scripting support enabled: the
rivemodule declaresWITH_RIVE_SCRIPTING=1andRIVE_LUAU=1and depends on the vendoredluau(its freshly generatedthirdparty/luau/luau.cppamalgamates the VM sources) andlibhydrogen(HYDRO_SIGN_VERIFY_ONLY=1) modules. Without the define,rive::File::read()has no importer forScriptAsset/ShaderAsset, so an in-bandFileAssetContentsbelonging to a script asset is routed to the previous asset's importer and tripsassert(!m_content)inFileAssetImporter::onFileAssetContents— which is what madetests/data/rive/viewmodel-lab.rivabort on load (and would have silently overwritten a font/image's contents in a release build). Scripts still only reach the VM when their in-band signature verifies, so unverified script assets load as inert content. - The vendored Rive runtime no longer prints
ScriptAsset doesn't have a generator function …once per scripted object per frame. A non-tools runtime registers only scripts whose in-band signature verified, so the generator lookup for an unverified script is expected to fail rather than worth a diagnostic. The message is dropped through a newpatchesentry on therivedependency intools/rive_update_manifest.json, so the nextjust rive_updatereapplies it instead of the noise coming back. - Android: the vendored Rive runtime no longer defines
RIVE_DESKTOP_GLon GLES platforms. It was defined wheneverYUP_RIVE_USE_OPENGLwas on andRIVE_WEBGLoff, so Android took Rive's desktop GL path andgles3.hppincluded glad'sgles2.h. glad'sglXxxfunction macros then aliased Rive's own (non-glad) extension entry points onto glad's identically named pointers, and the link failed withduplicate symbol: glad_glDrawArraysInstancedBaseInstanceEXTand eight siblings; had it linked, those pointers would also have been resolved by glad rather than byLoadAndValidateGLESExtensions(). Android now uses<GLES3/*.h>as intended. - Emscripten/WebGL: the OpenGL compute backend is no longer compiled, since WebGL (and WebGPU on the web) only reach GLES 3.0 and have no compute entry points. A new
YUP_RHI_USE_GL_COMPUTEconfig guards the GL compute pipeline, pass, factories and dispatch sites, replacing theYUP_RIVE_USE_OPENGL || YUP_LINUX || YUP_ANDROIDcondition that held on every Emscripten build and failed withuse of undeclared identifier 'GL_COMPUTE_SHADER','glDispatchCompute'and theGL_*_BARRIER_BITenums. - Android CI:
build_android.ymlnow tellsandroid-actions/setup-androidto install onlyplatform-tools, avoiding its removed defaulttoolspackage so the configure job can set up the SDK again
- The
YdspOptimizerTestssuite (tests/yup_dsp_jit/yup_YdspOptimizerTests.cpp) is enabled and now exercises each optimizer pass - constant folding (including loop-carried induction registers, which must not fold), algebraic simplification, copy propagation, dead-code elimination and loop-invariant code motion - in addition to the existing IR-lowering and execution-report checks. The individual passes are exposed onYdspOptimizerso tests can drive them directly; the four-pass loop (constant folding, algebraic simplification, copy propagation, DCE) is wired intoYdspOptimizer::runPasses, while LICM is exercised by its tests directly. - The AU and AUv3 wrapper tests no longer describe a stereo buffer list with a stack-allocated
AudioBufferList. Only the firstAudioBufferis reserved inside the struct, soAUStateTests.RenderProducesOutputand the twoAUv3BypassRenderTestsrender tests wrote past the object while fillingmBuffers[1].mDataByteSize, aborting the suite under AddressSanitizer with a stack-buffer-overflow. All three now build their lists with a newtests/yup_audio_plugin_client/yup_TestAudioBufferList.hhelper, which owns an allocation sized for the number of buffers asked for - The
yup_eventsPython tests no longer depend on a single fixed-duration pump of the message loop.next(juce_app)runs the dispatch loop for 20ms and returns, but a dispatched callback still has to re-acquire the GIL before the Python side runs, so on a loaded machine it can land after the pump has already returned -test_MessageListener::test_construct_and_postfailed this way on CI. The 22 call sites with a positive expectation now use a newpump_until(app, predicate)helper inpython/tests/utilities.py, which pumps in short slices until the condition holds or a 5s timeout expires. The four sites that assert a negative after pumping ("still zero because it was cancelled") deliberately keep the fixed pump, since polling a negative predicate returns immediately and proves nothing Componentnow befriends a singleComponentTestHelper<T>class template instead of accumulating one friend class per test suite; unit tests specialize it (e.g.ComponentTestHelper<Component>,ComponentTestHelper<ComponentEffect>) to reach private state.- Expanded the Rive viewmodel data-binding tests to run against
tests/data/rive/responsive-sliders.riv, the data-binding showcase fixture: concrete schema assertions for both its ViewModels (Slider_instance,Main) and their authored instances, nested viewmodel / dotted-path value access and color round trips, cross-handle shared-value visibility, property-change callbacks reporting dotted paths, and fullArtboardbind/advance/unbind coverage over the state-machine-driven "Main" artboard. - Added
tests/data/rive/viewmodel-lab.rivand the coverage that goes with it. The fixture ships four ViewModels in a known order (Details,Row,Panel,Lab), authored instances for each, and aLabschema declaring every property type the API models (string, number, boolean, color, custom enum, trigger, nested viewmodel, list, asset image, artboard reference). The new suites assert the schema contents (property order and types, the enum'sIdle/Running/Failedoptions, theRowschema's input flags), the authored values of theDefaultandPresetinstances - including both nested viewmodels, the three authored list rows and the per-instance list sizes - and the accessor contract for properties that carry no value (triggers, containers, symbol list indexes), plus bind/advance/unbind coverage of the "ViewModel Lab" artboard while writing values and mutating the row list. Every expected value is whattools/rive_inspect.pyreports for the fixture, which is now also part of the cross-fixture schema sweep. Because the fixture shipsScriptAssets, the suites skip with an explicit message when Rive is built withoutWITH_RIVE_SCRIPTING. Running them found one defect: unbinding an instance whose list drives anArtboardComponentListleaves the parent artboard's Yoga tree referencing the rows that the bulk teardown dropped, so the next layout pass walks freed nodes and crashes - that sequence is captured byDISABLED_UnbindingAndAdvancingKeepsTheArtboardUsablewith a pointer to the cause rather than weakened to match the current behaviour. - Nine test files were globbed into the IDE project but never
#included in their module's unity translation unit, so 141 tests had never been compiled or run since being written:yup_CodeEditorScheme.cpp,yup_ListBoxItem.cppandyup_PaintProfileStats.cpp(yup_gui),yup_Memory.cppandyup_TypeErasedObject.cpp(yup_core),yup_GraphicsContext.cppandyup_ImageFormatMetadataExtended.cpp(yup_graphics),yup_AudioDeviceManagerWindow.cpp(yup_audio_gui) andyup_AudioPluginLV2Format.cpp(yup_audio_plugin_host). All are now wired up, which took three kinds of repair: two defined a file-scope helper that a sibling in the same unity build already defined (makeSample,loadFromBlock), so the newly enabled copies are renamed rather than the working ones;yup_ListBoxItem.cpphad drifted against the graphics API (PixelFormatis no longer nested inImageand has noARGB, theImageconstructor now takes(w, h, format),DrawablePathfolded intoDrawable, andStringhas no(count, char)constructor); andyup_AudioPluginLV2Format.cppis now wrapped in#if YUP_AUDIO_PLUGIN_HOST_ENABLE_LV2, mirroring the guard the module puts aroundLV2Format, since the test target does not enable LV2. Wiring them up surfaced two genuine defects the dead tests had been written against: the PNG raw-chunk writer bug below, andListBoxItem::setIconbeing an unimplemented stub - the four tests depending on it are markedDISABLED_with a pointer to the TODO rather than weakened to match the stub
- YDSP state is now segmented as
[scalars][arrays]and the kernel ABI carries both base pointers (state+ newstateArraysinYdspKernelContext/YdspEventContext, withYdspCodegen::stateScalarSizereporting the split). Every scalar slot (including the ring write-pointers of the@delay primitives) lives in the scalar segment head, and all arrays follow in their own segment. Previously, scalars were addressed at byte offsets that grew with array state (e.g. 50476 for a reverb kernel), which AArch64 could not encode as an LDR/STR immediate and failed to assemble withInvalidDisplacement(ldur w2, [x3, 50476]). Array state can now grow arbitrarily (delay lines, reverb rings) without pushing scalar slots out of range; an out-of-range materialization fallback remains as defense in depth. The optimiser's per-type array-base shift pass is gone (array bases are per-type element indices into the array segment). - SDL windowing: partial repaints now grow the dirty area by half a pixel before rounding it out to whole pixels. Rive applies rectangular clips as anti-aliased coverage rather than a pixel-exact scissor, so geometry touching a component's clip edge could bleed a tiny coverage into the adjacent pixel row; with a preserved render target that row was never redrawn and the bleed accumulated into a persistent line just outside components that repaint continuously (visible around the
SpectrogramComponentat 1x scale). The extra border lets the parent repaint those pixels every frame. - AUv2 wrapper: an input bus that is not fed during a render cycle is now presented to the processor as a null-channel view.
buildInputBusViewsfilled the per-bus channel pointers only for the channels it actually received, leaving the rest at whatever the previous render had stored there, so a sidechain input the host stopped feeding (inactive element or a failedPullInput) kept pointing at that element's stale audio instead of reading as silent - contradicting the comment onpullAuxiliaryInputElementsand theAudioBusBufferView"null for an inactive or silent bus" contract. The input path now clears each bus's slots first, exactly asbuildOutputBusViewsalready did for outputs; the AUv3 wrapper never had the problem because it maps input views onto its own scratch buffers MessageManageron Apple platforms:runDispatchLoop()andrunDispatchLoopUntil()now wrap their loop body inYUP_TRY/YUP_CATCH_EXCEPTION, matching the generic implementations inyup_MessageManager.cppthat#if ! (YUP_APPLE || YUP_WASM)compiles out on these platforms. The.mmreplacements previously caught onlyNSException, a disjoint set fromstd::exception, so a C++ exception thrown by a message callback escaped the dispatch loop instead of reachingYUPApplicationBase::sendUnhandledException()— which madeunhandledException()unreachable on macOS and iOS, and killed the application on the first failure.- The SDL render thread now routes exceptions from
renderFrame()throughYUP_CATCH_EXCEPTIONas well. It never passes through a dispatch loop, so an exception escapingpaint()reachedThread::threadEntryPoint(), which only asserts — rendering then stopped permanently with no diagnostic in release builds. - PNG:
png/iCCPandpng/cHRMraw chunks are now actually written. The iCCP branch in the writer was an emptyifbody with a "for simplicity, write as unknown chunk" comment that never wrote anything, andpng/iCCPwas excluded from the unknown-chunk loop; cHRM was collected but silently dropped by libpng, because on write libpng only emits unknown chunks whose name marks them safe-to-copy (a lowercase fourth letter) unlesspng_set_keep_unknown_chunkssays otherwise -eXIfis safe-to-copy and survived,cHRMandiCCPare not and did not. Both are now registered withPNG_HANDLE_CHUNK_ALWAYS. Thepng_unknown_chunkis also value-initialised, since libpng copies all five name bytes and the terminator was left indeterminate FlexBox: thegapis no longer applied after the last item on each line. It was added unconditionally after every item, so when items grew or shrank to fill the container exactly (e.g.flexGrowitems withgap), the trailing gap pushed the final item past the container's main-axis edge and it got clipped (e.g. the last panel in each row of theLayoutexample).ShaderTranspiler: GLSL ES fragment output now defaults toprecision highp float;instead of SPIRV-Cross'sprecision mediump float;default. Desktop GLSL implies highp and glslang records no precision decorations for it, so SPIRV-Cross could only re-qualify declared variables; anything left to the fragment default — uniform block members and inlined expression intermediates — silently ran at mediump on OpenGL ES, corrupting fp32-exact math such as the fixed-point field codecs of the GPU fluid simulation demo (pixelated dye that never fades on Android). The ESSL default is now highp in both the emit and reflect paths.- Fixed data races in
KMeterState: the per-channel getters (getPeakLevel/getAverageLevel/getPeakHoldLevelwith an explicit channel index) read the plainChannelStatefloats the audio thread updates, bypassing the atomics the aggregate (channel -1) path uses — those levels are now atomic, with the audio-side ballistics computing on locals and publishing once per block. The runtime configuration scalars (scale,meteringStandard, fall/hold times, over threshold/mode) are now also atomic since the processing thread reads them while setters run on the UI thread, andsetMeteringStandard/setIntegrationTime/setPeakFallTimeno longer mutate the loudness filters and level processors from the calling thread — they set a flag thatprocessPendingAudio()applies on the processing thread that owns them. The class now usesstd::atomicthroughout (previouslyyup::Atomic) StyledTextcaret bounds, hit-testing and selection rectangles now use line-relative glyph x positions computed with the same accumulation as drawing, instead of rive's paragraph-relativeGlyphRun::xpos. Character positions were wrong on soft-wrapped lines (off by the width of all preceding text in the paragraph) and selection was drawn shifted on wrapped text; the caret at the first character of a wrapped line now lands on that line's left edge.- iOS applications now use the
UIScenelifecycle, removing UIKit's legacy lifecycle warning and ensuring SDL windows are created for the connected scene. - Offscreen GPU rendering now supports recursive targets on Metal, OpenGL/GLES, and D3D11, so Lottie alpha/luma mattes, isolated-opacity layers, and cached precomps retain GPU compositing when rendered into an
ImageorGpuCanvas. EachRenderableTargetleases a Rive render context exclusively for its lifetime and returns it to the pool when destroyed. Repeated Lottie matte and precomp renders now reuse their canvases rather than allocating GPU textures each frame. Metal child targets allocate only their Rive render-canvas output texture; the CPU readback staging texture is created only when pixels are requested. - Fixed undefined offscreen contents when nesting pooled render targets on all GPU backends. Render context slots were recycled whenever no frame was currently active, so two long-lived targets could share one slot; once their frames nested - which happens as Lottie matte and precomp layers cross their in/out points and the nesting order changes between frames - the inner target skipped
beginFrameand was then flushed against the outer target's frame descriptor. - Lottie: a matte layer no longer paints another matte layer's content. Drawing a matte result only queues a reference to its canvas texture, which the enclosing frame resolves at flush time, but the canvas lease was released as soon as the layer finished. Since every matte in a composition is sized to the same fitted rectangle, the pool handed the same canvas triple to the next matte layer, which overwrote the pixels already queued and left only the last matte visible (e.g.
world_locations.json's four matted dots collapsed to one and its continent outlines disappeared;insta_camera.jsonlost its animated circles). Leases are now held until the composition render completes. GpuFramenow waits for the GPU before releasing the texture views, uniform buffers and samplers it keeps alive for its encoded render passes. Those passes reference them by raw pointer, andsubmit()does not block, so letting a frame go out of scope freed them while the GPU was still reading — corrupting the pass output progressively, as the freed memory only starts being handed back out after the allocator has churned for a while (the growing magenta flashes inbell.json).waitForGPU()is now only needed explicitly when results are required before the end of the frame's scope, and is idempotent so waiting explicitly costs no more than one stall. Move-assignment drains the frame it replaces for the same reason.AffineTransform::getScaleFactor()is now independent of rotation. It averaged the absolute values of the matrix diagonal and ignored the shear terms, so a rotated transform reportedscale * cos(angle)— falling to zero at 90 degrees. It now measures the lengths of the transformed basis vectors. Lottie precomposition and matte canvases are sized from this value, so a layer under an animated rotation (e.g.bell.json, whose precomposition is parented to a rotating null) requested a different pixel size on every frame, reallocating its canvases mid-frame and flashing while a queued draw still referenced the previous ones.- Lottie: the matte canvas pool now replaces an idle slot of a different size instead of appending a new one. Nothing removed slots, so a layer whose on-screen size changed every frame added three canvases — each leasing a Rive render context — per frame, without bound.
GpuCanvas::create()takes astd::optional<Color> clearColor, defaulting to transparent black, and fills the new canvas with it so it is safe to sample before anything is drawn into it (passstd::nulloptto leave the contents undefined). The backing texture is allocated uninitialized and a 2D frame whose draw list ends up empty is not guaranteed to honour itsloadAction=clear, so a canvas could previously be composited while still holding undefined GPU memory (the magenta flashes inbell.json, whose only content is one matted precomposition). The clear is issued through the newGpuDevice::clearOffscreen(), which encodes it with the backend's native API — a clear binds no pipeline, buffers or samplers, so it needs neither a render pass nor a submit/wait cycle.Artboard::clear()andupdateSceneFromFile()destroyed the rive artboard before the scene (StateMachineInstance) that references it, so destroying the scene calledcleanupFocusTree()on freed memory (ASAN use-after-free). The scene is now reset before the artboard.ArtboardViewModelInstancepath resolution downcast a property torive::ViewModelInstanceListwithout checking its type whenever a path segment was an index, so a path like"score.0"(wherescoreis a number) asserted in debug builds and read a garbage vector in release ones. Reachable from every path-taking accessor. The four uncheckedViewModelPropertyEnumdowncasts behind the enum accessors were guarded the same way.Artboardonly drained its state machine's reported events frommouseDrag(), so the documentedonPropertyChanged/propertyChangedcallbacks never fired for events reported by an advance, a pointer press or a transition — Rive clears the queue at the top of the next advance, and a resize or abindViewModelInstance()(both of which advance) silently swallowed the frame's events. Every advance now drains through one helper, and every pointer handler drains after itspointer*call.Artboard::notifyNodeBoundsChanged()iterated a member scratch array while invoking user callbacks that are allowed to callsetLayout()/setAlignment(), which re-entered the function and cleared the storage the outer loop was walking (and theconst String&its callee held). Nested passes are now skipped, and the loop tolerates a callback replacing or unloading the file.ArtboardViewModelInstanceinvoked its property-changed callback in place, so a callback that re-armed or cleared itself (a one-shot listener) destroyed the closure that was executing — something the header explicitly documents as supported. The dispatch now runs a local copy.Artboard::applyNodeAttachment()took its attachment record by reference straight out of theattachedComponentsmap, then readoptions.applyTransformfrom it afterComponent::setBounds(), which firesresized()synchronously. Aresized()that detached the component (or attached it to another node) destroyed that map entry mid-call. The record is now taken by value.Artboard::setFile (nullptr)dereferenced the null file instead of unloading, and the file-taking constructor did not callsetOpaque (true)while the other one did.ArtboardFile::AssetInfo::uniquePathwas aFileholdingrive::FileAsset::uniqueFilename(), which is a bare file name ("logo-1234.png") and not an absolute path — so every asset-resolving load hitjassertfalseinFile::parseAbsolutePath()and then silently resolved the name against the current working directory. Renamed touniqueFilenameand retyped as aString; resolve it against your own asset directory withFile::getChildFile().ArtboardFile::load()ignored the result ofreadIntoMemoryBlock(), so a stream that yielded nothing was reported as"Malformed artboard file"rather than as a read failure.GpuCanvas::beginDraw()now drops the target's cachedGpuTexturewrap, as its documentation already claimed. The wrap memoizes the Rive texture handle it resolved, so a pooled canvas reused across frames kept handing out the handle resolved on the frame it was first sampled.- Lottie: a failed matte composite no longer blits undefined GPU memory over the matted layer. The result canvas is written only by the composite render pass - nothing else clears it, and its backing texture is allocated uninitialized - but the pass result was ignored and the texture composited regardless, flashing an arbitrary color. The renderer now falls back to the geometric-clip matte path when the composite fails.
- Lottie: a paint-less nested group now contributes its geometry to the enclosing group's paints with its own modifiers applied. The geometry was rebuilt from raw shapes, dropping the nested group's trim, repeater, merge-paths and rounded-corner modifiers, which is what defines the outline: RubberHose rigs draw a limb as a 4-point star trimmed to a quarter, so the parent stroke painted the whole star instead of an arc (the stray stars in
mughead.jsonandpumped_up.json). - Lottie: track mattes (alpha, alpha-inverted, luma, luma-inverted) now composite the matte source's rendered alpha - including its fill opacity, gradients, and anti-aliased edges - instead of hard-clipping the target to the source silhouette. The matte source and target are rendered into offscreen GPU buffers (sized to the fitted on-screen resolution) and multiplied by a fullscreen matte-composite shader. A partially transparent matte source now shows through correctly (e.g.
matte_two_item_with_lowerlayer.json, whose 65%-opacity source blends the white matted ellipse to pink over the red layer beneath). Falls back to the previous geometric-clip behaviour when no GPU is available (e.g. headless rendering). - Lottie:
EllipseShapepaths now start at the top (12 o'clock) and follow the shape direction (clockwise ford == 1, counter-clockwise ford == 3), matching Lottie's convention. Previously they started at the right (3 o'clock) going counter-clockwise, which placed trimmed arcs at the wrong position (e.g. the expanding rings inworld_locations.jsonwere cut short on the right). Path::withRoundedCorners()left one corner sharp on closed subpaths whose geometry ended with an explicit segment back to the start vertex (as produced by Lottie beziertoPath()). The duplicated start/end point formed a zero-length edge that made that corner degenerate. The trailing duplicate is now dropped, and corners are rounded with a cubic arc (circle kappa) instead of a single quadratic through the vertex, so a square with a full Round Corners modifier becomes a proper circle (e.g. the morphing loader shape inloader.json).- Lottie: trailing top-level modifiers (trim, repeater, rounded-corner) in a shape layer now apply to every preceding top-level group in the run, not just the last one, so a single trim animates all shapes it should (e.g. the knife in
it's_lunch_time!.json, and the segmented strokes inimprint.json/fingerprint_success.json). Trailing paints similarly reach all preceding paint-less groups. - Lottie: animated properties driven by an AfterEffects
loopOut('cycle')expression (AnimationProperty<T>::LoopMode) now repeat their keyframe range instead of freezing on the final value once playback passes the last keyframe. Fixes pulsing markers vanishing after their first cycle (e.g. the orange location circles inworld_locations.json). - Lottie: precomposition layers are now rasterized to an offscreen texture sized to the on-screen device resolution instead of the fixed composition size, so precomps no longer look blurry when the animation is scaled up (e.g.
tractor.json). - Lottie: layers with partial (animated) opacity are isolated into a transparency layer for correct compositing; this offscreen buffer is now sized to the fitted on-screen resolution instead of the composition size, so small compositions no longer look blurry when scaled up (e.g.
spin,_lil_loader_v2.json, a 90x90 composition whose fading "stick" layers were rasterized at 90px and upscaled). - Lottie: Merge Paths (
mm) is now supported (AnimationMergePaths). Boolean modes (Add/Union, Subtract, Intersect, Exclude) combine the preceding path geometry with the matching boolean operation, while the plain "Merge" mode concatenates paths and lets the fill winding rule form counters (holes). Nested paint-less groups only feed their geometry to the parent group's fills/strokes when a Merge Paths modifier is present; otherwise nested groups stay self-contained so paint-less construction guides are not accidentally filled (fixes stray star/cross shapes and per-frame overhead inpumped_up.jsonandmughead.json). Fixes shapes built from merged sub-paths rendering only partially (e.g. the red windmill sails inwindmill.json) without filling in letter counters (e.g. the holes in "O"/"A" ingoal.json). - Lottie: the AfterEffects inertial-bounce ("overshoot") position expression (
amp/freq/decay) is now approximated viaAnimationTransformInertialBounceParams, producing the decaying oscillation past the last position keyframe. Fixes elements that dropped in without the expected bounce (e.g.windmill.json). - Lottie /
AnimationRenderer::renderComposition: content that extends beyond the composition viewport (e.g. shapes with coordinates outside thew/hbounds, as injolly_walker.json) now clips to the fitted composition rectangle instead of the full target bounds, so it no longer spills into the letterbox / pillarbox area when the target rectangle is not the composition's aspect ratio. AnimationTransform::positionAt()spatial bezier motion paths were nearly straight instead of curved: the second control point used the next keyframe's incoming tangent (k1.tangentIn) rather than the current segment's own tangent (k0.tangentIn). In Lottie bothtoandtibelong to the keyframe starting the segment, so a circular motion path (e.g. a shape orbiting on a bezier arc) collapsed toward linear interpolation.- OpenGL / WebGL: the main frame's rive flush went silently blank (draws degenerate, screen frozen on the last good frame) whenever a
GpuCanvascommitted mid-frame.endOffscreen()'sunbindGLInternalResources()wipes the shared GL texture units, but the main render context's internal textures (tessellation/gradient/feather/atlas) were only rebound atbegin()- beforepaint()- so any offscreen 2D flush during paint left the main flush sampling incomplete textures (no GL error; GLES returns zeros). The GL backend now callsinvalidateGLState()on the flushing context immediately before everyflush()(main frame and offscreen), making each flush self-contained regardless of how many rive/ore contexts interleave on the one real GL context. FixesSpinningCubeDemoon WASM/WebGL2 appearing frozen (with sporadic 5-15 s updates) and the page turning sluggish while the app still reported ~57 FPS. Graphics::drawTexture/drawImage/ transparency layers rendered nothing (transparent) whenever the rive frame ran in atomic interlock mode - always the case on the iOS simulator, and on any platform when raster ordering is disabled. The composite was implemented as a path draw with an image paint, which atomic-mode shaders cannot sample;Graphics::renderTexturenow routes throughrive::Renderer::drawImage, which falls back to a dedicated image-rect draw in atomic mode. Fixes invisible Lottie precomps/mattes,GpuCanvascomposites, and the SpinningCube demo output on the iOS simulator.- OpenGL / WebGL:
GpuCanvastextures drawn withGraphics::drawTexture(Lottie precomp caches and matte composites) rendered vertically flipped, because the GL canvas source texture is stored bottom-up.GpuTexture::getOrAdoptGpuTexture()now prefers the Y-flipped sampled mirror - kept fresh at each canvas flush - matching whatGpuRenderPassalready did for sampled inputs. No change on Metal/D3D, where the mirror is null. - SDL3 windowing: mouse move/drag was broken on touch platforms (iOS, Android). Motion was synthesized only by polling
SDL_GetGlobalMouseState, which has no backend implementation there and falls back to window-relative coordinates, so subtracting the window position shifted every move. Touch platforms now consume the touch-synthesizedSDL_EVENT_MOUSE_MOTIONevents directly; desktop keeps the global-cursor poll (needed for embedded plugin editors). - SDL3 windowing: mouse drag events were lost inside embedded plugin editors (notably on macOS, where the host owns the native application so SDL never receives Cocoa mouse focus and suppresses drag motion). Dragging is now synthesized by polling the global cursor while a button is held, on the message thread, for all platforms.
Slidercould get stuck showing its hover color after a touch drag:mouseEnter/mouseExitnever fire for touch (no hover phase, and drag capture bypasses them for the mouse too), so releasing outside the slider's bounds left the hover state on.mouseUpnow clears it directly when the pointer isn't over the slider anymore.- UBSAN and ASAN fixes throughout the codebase
- AUv3 plugin host bypass is now connected to the processor: the wrapper-owned bypass parameter is created and drives
processBlockBypassed, and host bypass state is persisted/restored inside theYUPProcessorStateblob (legacy raw processor state still loads) - Added bypass parameter handling tests for the AU, CLAP, and VST3 plugin client wrappers (routing to
processBlockBypassed, bypass state round-trip, and text/value conversion) - Windows toasts emit the
scenarioattribute with the spellings the toast schema declares (reminder/alarm/incomingCall) rather than the capitalised WinToast ones, which are not part of the enumeration. Schema conformance only — it is not the cause of the toasts that fail to display on Windows 11, seedocs/Windows Toast 80070490 Analysis.md - Windows toasts report a real permission state instead of always claiming
granted:ToastNotification::getPermissionState()/requestPermission()now queryIToastNotifier::get_Setting(), so an application, user, group policy or manifest level block is visible to the caller. The setting is also logged next to the payload. Note that it does not cover Do Not Disturb or the per-app "show notification banners" switch, which suppress the on-screen banner while still delivering the toast to the notification center - Windows toasts no longer hand
put_ExpirationTimea stack object that dies at the end of the enclosingifblock. The notification retains thatIReference<DateTime>for its whole life, so it was already dangling by the timeShow()read it; it is now a reference-countedComBaseClassHelperthat the notification keeps alive TypeErasedObjectnow relocates its payload through the payload's move constructor instead of a byte copy. Moving a payload that points into itself - libstdc++'s small-stringstd::string, astd::mapheader, or anything caching a member's address - left those pointers aimed at the dead source buffer, soGpuPipeline::Impl's entry-point strings freed a stale stack address when the pipeline was destroyed. glibc reported it asfree(): invalid pointerand killedyup_testsinGpuAttachmentMockTestson Linux, while libc++'sstd::stringhas no self-pointer, so macOS never saw it
- Added a dedicated DSP documentation area (
docs/dsp/) coveringyup_dspend to end: math/windowing/noise, FFTs and spectral analysis, filter design and filter implementations, dynamics and metering, onset detection, convolution and delay, resampling, and time-stretching/pitch-shifting docs/dsp/yup-dsp-language.md(the YDSP language reference) is now linked fromdocs/dsp/index.md's toctree - it previously built but was unreachable from the docs site. Its §2.7 EBNF now covers the bitwise operators (& | ^ ~ << >>, at their actual precedence, which is tighter than comparisons unlike C) that were already implemented but undocumented; §2.8 lists the previously-undocumentedasinh/acosh/atanh/round/copysignintrinsics and the new integer overload ofmin/max/clamp/abs/sign(including theabs(INT_MIN)edge case); and §2.7/§3.2 each gain a sentence clarifying that unary~(bitwise not) and the graph algebra's binary~(recursion) are unrelated operators in separate grammars, not an overload of one operator.
- Android window support with
YupActivityJava class (#29, #34) - Java bytecode compilation via
yup_android_java.cmake(#53) - External storage permissions (
READ_EXTERNAL_STORAGE/WRITE_EXTERNAL_STORAGE) for file access (#61)
- iOS CI pipeline with Xcode toolchain (#8)
- Updated minimum deployment targets: iOS/tvOS 13.0, watchOS 6.0, macOS 11.0 (#72)
- ARC enabled by default on Apple platforms (#91)
- iOS Simulator-specific framework groups (
iosSimFrameworks/iosSimWeakFrameworks) in module declarations (#48)
- macOS message loop reworked: time-sliced event dispatch via
CFRunLoopRunInModetargeting ~60 Hz; quit event registered withNSAppleEventManagerfor proper Apple Event quit handling (#47) NSSupportsSuddenTermination = falseadded to macOSInfo.plist(#47)
- Full Emscripten/WASM support including
AudioWorkletaudio device (#25) - WASM threading with exported runtime methods (#61)
-msimd128compile flag and configurable stack size for Emscripten targets (#98)
- SDL2 integration with libpng, libwebp, and rive_decoders (#37)
- Improved SDL/JUCE Message Manager dispatch loop (#38)
- SDL symbol namespacing to prevent linker conflicts in Apple platform plugins (#112)
- Reworked rendering backend selection API:
YUP_RIVE_USE_D3D,YUP_RIVE_USE_METAL,YUP_RIVE_USE_OPENGL,YUP_RIVE_USE_DAWN(#24) - Headless graphics context and no-op Rive factory for offscreen rendering (#32, #52)
- SVG rendering support (#56, #64)
- SVG 1.1 spec compliance: blend modes, patterns, polygon/polyline (#100, #118)
- Path API improvements with comprehensive examples (basic shapes, arcs, curves, transforms, advanced) (#56)
createStrokePolygon()with feather effects (#55)- Improved color management and gradient editor (#87)
- Improved CPU and GPU image rendering (#39, #87)
Color::brighter()/darker()madeconst(#32)AffineTransform::inverted()constexpr method (#19)- constexpr math utilities:
juce_abs(),jmap(),jlimit(),findMinimum()/findMaximum(),nextPowerOfTwo(), and more (#18) ColorGradient::Spreadenum:Pad,Repeat,Reflecttiling modes withwithSpread()builder (#119)CubicBezierclass:pointAt(),derivative(),length(),splitAt(), bounding box, and intersection (#119)
- New image format I/O framework:
ImageFormat,ImageFormatReader,ImageFormatWriter,ImageFormatManager- plugin-style registry with magic-byte detection and animated image support (#119) - BMP image format: reader (1/4/8/16/24/32-bpp, RLE4/RLE8, palette) and writer (24-bpp uncompressed), controlled by
YUP_IMAGE_FORMAT_BMP(#119) - PPM/PGM/PBM (Netpbm) image format: full P1–P6 plain and binary read/write, controlled by
YUP_IMAGE_FORMAT_PPM(#119) - PNG image format via
libpng: grayscale, grayscale+alpha, RGB, and RGBA at 8- and 16-bit depths, controlled byYUP_IMAGE_FORMAT_PNG(#119) - JPEG image format via
libjpeg: quality-level encoding, controlled byYUP_IMAGE_FORMAT_JPEG(#119) - WebP image format via
libwebp, controlled byYUP_IMAGE_FORMAT_WEBP(#119) - Animated GIF image format via
libgif: per-frame delay, loop count, animated write API (beginAnimation/writeFrame/endAnimation), controlled byYUP_IMAGE_FORMAT_GIF(#119) Image::loadFromData()reimplemented viaImageFormatManager(#119)
GraphicsContext::OffscreenTargetabstract interface for opaque platform GPU offscreen resources (#119)- Offscreen API on
GraphicsContext:createOffscreenTarget(),beginOffscreen(),endOffscreen(),readOffscreenPixels()- implemented for Metal, OpenGL, and D3D backends (#119) Graphicsconstructors for rendering to anImageorOffscreenTargetoutside the main frame cycle (#119)ImagegainedrenderCanvasbacking (RenderCanvas) alongside texture for offscreen render-to-texture;duplicate()re-enabled with proper deep copy (#119)Graphics::TransparencyLayerRAII class for isolated group opacity compositing: renders into an offscreen target and composites back at the given opacity oncommit()(#119)
- New
yup_animationmodule: Lottie-compatible animation engine depending onyup_coreandyup_graphics(#119) Animation: high-level handle withloadFromFile(),loadFromData(),loadFromStream(),renderFrame(),renderAtTime(),renderAtProgress(),toJson(), andsaveToFile()(#119)AnimationPlayer: stateful playback controller with forward, reverse, and ping-pong direction modes, looping, variable speed, frame-range clamping, seek, andonFrameChanged/onLoopCompleted/onPlaybackEndedcallbacks (#119)AnimationEasing: cubic bezier easing with named presets (linear,easeIn,easeOut,easeInOut,hold) andfromLottieTangents()import (#119)AnimationProperty<T>: generic animated property with keyframe interpolation; specializations forfloat,Point<float>,Size<float>, andColor(#119)AnimationTransform: animated anchor, position, scale, rotation, and opacity with conversion toAffineTransformat a given frame (#119)- Full animation data model:
AnimationComposition,AnimationGroup,AnimationLayer,ShapeLayer, shape types (ellipse, rect, path, star, merge, trim, repeater, polystar), paint types (fill, stroke, linear/radial gradient), and modifiers (#119) LottieReader: parse.jsonand.lottie(ZIP) files from file, string, or stream into the animation data model (#119)LottieWriter: serialize the animation data model back to Lottie JSON (pretty or compact) with full round-trip support (#119)LottieExpressionEvaluator: JavaScript expression evaluator for Lottie property expressions viaJavascriptEngine(#119)AnimationRenderer: renders anAnimationCompositionto aGraphicscontext - layer hierarchy, parent-child transforms, matte layers (track-matte), shape fills/strokes/gradients, and image layers (#119)AnimationFrameExporter: exports individual or all frames toImageobjects via offscreen GPU, and exports animations to animated GIF files (#119)AnimationRendererno longer renders a precomposition into an offscreen GPU target unless more than one layer shares that asset: a single-reference precomp now draws straight into its parent, removing a render target, its clear and a full GPU flush per nesting level. Nested precomps previously each opened their own target- Fixed
AnimationRendererrebuilding a precomposition's viewport clip with a path boolean op on every layer: the renderer already intersects two rectangular clips itself (in float space, through its clip-rect path), and the boolean op turned the viewport into a polygon with a redundant vertex per crossing that then had to be re-tessellated as a clip path. Clips that cannot overlap the one in effect now cull the layer or precomp outright - Fixed the per-layer mask clip cache never being reusable while playing back: it was keyed by frame number even for masks that never animate, so a static mask re-ran its boolean ops on every frame. A static mask is now cached on the layer itself (keyed by composition size, which the mask bounds derive from), and an animated one still caches per frame
AnimationRenderer's parent-transform resolution is now linear: every layer's accumulated transform is resolved once and memoized, instead of rescanning the whole layer list (and re-testing every unresolved parent) until nothing more resolves. A cyclic parent chain is bounded rather than retriedAnimationRendererno longer isolates a layer behind an offscreen composite when that cannot change the result: a solid, image, text or null layer at partial opacity already folds its opacity into its single paint, so compositing it through a render target reproduced the same pixels at the cost of a full render target, a clear and a GPU flush per layer per frame. Layer types that draw several overlapping primitives, and any layer carrying a drop shadow, still composite offscreen- Fixed
AnimationRendererisolating a layer over a full-composition offscreen target: the target is now sized to the layer's own content box (intersected with the composition viewport) and the content shifted to match, so a small layer in a large composition no longer pays for the whole composition in clear, flush upload and composite. Rasterization resolution is unchanged - Opacities of 0.999 and above are now treated as fully opaque. Exporters write values like 99 or 99.9 for layers meant to be seen at full strength, and isolating one behind an offscreen composite to reproduce a sub-1% difference in alpha cost more than the difference was worth
- Fixed
AnimationRenderertreating a precomp as shared when its referencing layers are never on screen together: reference counting now counts only the layers the frame actually draws (hidden, matte-source and out-of-range layers excluded, each nested level evaluated at its own frame). A composition split into sequential time slices - a common export shape - is therefore drawn like any single-reference asset instead of through a full-size offscreen target on every frame
- VST3 plugin support (#44)
- Audio Unit (AUv2) plugin support (#93, #106)
- Barebone Audio Unit (AUv3) plugin support (#122)
- Barebone AAX plugin support (#122)
- Barebone LV2 plugin support (#122)
- Audio Plugin Host for AUv2, VST3, and CLAP (#93, #98, #106)
- Standalone plugin support with improved audio parameters (#46)
- CLAP/VST3/AU validators and code signing (
YUP_ENABLE_VST3_VALIDATOR, etc.) (#106) - pluginval integration for automated VST3 validation (#67)
- Sidechain and multi-bus audio input support across VST3, CLAP, AUv3, AUv2, AAX, and LV2:
AudioBusgains aRole(Main/Auxiliary) andisDefaultActive,AudioProcessContextexposes per-businputs/outputsviews (AudioBusBufferView) withgetMainInput()/getAuxiliaryInput()/getMainOutput()accessors, and secondary input buses are forwarded to the processor instead of being discarded
- New
yup_audio_formatsmodule:AudioFormat,AudioFormatManager,AudioFormatReader,AudioFormatWriter, WAV codec (#51) - Opus, MP3, FLAC, AAC, CoreAudio, and WMF codec support (#86, #88)
- New
yup_dspmodule with FFT/windowing via ooura, pffft, vDSP, IPP, FFTW3 (#71) - Basic IIR filter implementations (#71)
- Linkwitz-Riley crossover filters (#71)
- FIR filter (#75)
- Partitioned convolution (#75)
- Oversampler and Resampler (2×/4×/8×) (#97)
- Noise generators (#71)
- Virtual analog filters:
AnalogTwoPoleFilter,AnalogVowelFilter,AnalogKorg35Filter,AnalogMoogLadderFilter,AnalogRolandDiodeFilter,CombFilter(#103) - Spectral processor (#116)
- Onset detectors (
SpectralFluxandComplexFluxODF) with perceptual filter bank (#117) - Time-domain and Frequency-domain stretching with backend selection (homebrew PSOLA plus bungee) (#104)
- Distortion processors with oversampling:
TanhDistortionProcessor,BlunterSoftClipperProcessor,AaIirHardClipperProcessor(#108) - Click-less fractionally addressed delay (FAD) (#109)
- Emscripten
AudioWorkletaudio device (#25) - MIDI 2.0 / Universal MIDI Packets (UMP) implementation (#83)
KMeterStatede-interleaving buffers pre-allocated inprepare(), eliminating per-block heap allocation in the audio callback (#119)
SynthesiserVoiceconverted toReferenceCountedObjectwithPtr = ReferenceCountedObjectPtr<SynthesiserVoice>;Synthesiser::addVoice()now acceptsSynthesiserVoice::Ptr(#82)
- New
yup_audio_graphandyup_audio_plugin_hostmodules (#93) - Thread-safe
BufferingAudioSourcewith atomicnextPlayPos(#98) AudioPlayHead::getContinuousTimeInSamples(): continuous sample time without loop-reset (#35)
- Customizable theming/skinning system (#13)
mouseDoubleClick()virtual method with configurable threshold (#14)KeyboardFocusModeenum replacing booleansetWantsKeyboardFocus(); addedtextInput()callback (#16)PopupMenuandComboBoxcomponents (#57, #62)- Native file chooser via
FileChooser(#61) - Component paint profiling:
PaintProfilerwith ring-buffer stats (min/max/mean/p50/p95/p99) (#95) ComponentNative::getGraphicsContext()virtual method allowing components to access the GPU context for offscreen operations (#119)MouseListenerweak-referenceable interface for all mouse events;Component::addMouseListener()/removeMouseListener()(#30)- Improved slider components (knob, linear, range) and button components (#70)
- Unified drag-and-drop support in
Component:isInterestedInDrag()/itemsDropped()virtuals with a fluentDragAndDropDatapayload (files and text on SDL, URIs reserved for future backends); drops dispatch to the topmost interested component and bubble up to parents - Added unit coverage for
SystemClipboarddata formats andComponentdrag-and-drop callbacks - Safe area support:
Component::getSafeAreaBounds()andsafeAreaChanged()virtual (backed byComponentNative::getSafeAreaBounds()), so content can avoid display cutouts and system bars on mobile devices - High dpi support on Windows and Linux X11: window bounds, screen geometry and input coordinates are now logical points everywhere (converted at the SDL boundary), so windows and content scale with the display scale like on macOS; live display scale changes resize the native window keeping the logical size
TextEditorandLabelcomponents (#16, #55)- Improved fonts: better layouting, variable font axis manipulation, embedded fallback font (#55)
- Clipboard support: text, MIME-typed data with lazy callbacks, and primary selection (#55)
- New
yup_audio_guimodule (#70) - MIDI keyboard component (#70)
- Filter frequency response visualisation (#71)
- Spectrum analyser component (#71)
- Spectrogram component with peak/RMS/power/PSD level modes (#102)
- Oscilloscope and spectrum analyzer display processors (#109)
- New
yup_data_modelmodule (#15) UndoManagerwithTransaction,ScopedTransaction,UndoableAction(#15)DataTreehierarchical data structure with builder pattern and transactional mutations (#73, #74)DataTreequery support (#74)DataTreeschema validation (#74)CachedValue<T>for type-safeDataTreeproperty references (#74)DataTreecompleteUndoableActionsuite for transactional mutations:PropertySetAction,PropertyRemoveAction,RemoveAllPropertiesAction,AddChildAction,RemoveChildAction,RemoveAllChildrenAction,MoveChildAction,CompoundAction- all with full undo/redo semantics (#82)Identifierusable asstd::unordered_mapkey viastd::hashspecialization (#27)
- Improved artboard placement (#17, #43)
- Shared Rive file across multiple
Artboardcomponents (#43) - State machine inputs:
setNumberInput(),setBoolInput(),triggerInput()(#17, #43) advanceAndApply()anddurationSeconds()for timeline control (#17)- State machine event handling (#43)
- Multi-artboard component support (#43)
constructAt()/destroyAt()/voidify()inmemory/yup_Memory.h: portable replacements forstd::construct_at/std::destroy_at, used byTypeErasedObjectTypeErasedObjectnow supports class template argument deduction (deduction guide sizes storage to the stored value) and move construction / assignment from a smaller-sizedTypeErasedObjectSqliteDatabasewithStatementandTransaction(#94)- Perfetto profiling:
YUP_ENABLE_PROFILING,Profilersingleton,YUP_PROFILE_START/YUP_PROFILE_STOPmacros (#20) Watchdogfile watching utility (#50)URLcopy and move constructors (#58)messageThreadIDmade atomic (removed mutex) (#26)ReferenceCountedObject::incReferenceCount()/decReferenceCount()madeconst(#28)DatagramSocketmulticast overloads with local IP (#60)ResultValue<T>::valueOr()(#96)AudioSampleBuffer::fill()overloads (#110)JavascriptEngine::executeWithResult()to execute a code block and capture the last expression result (#119)JavascriptEngine::registerNativeFunction()for top-level native function registration by name (#119)- JavaScript
$accepted as valid identifier character, required for Lottie expression compatibility (#119) - WASM: POSIX file API extended with
symlink(),dirent.h,fnmatch.h,utime.hsupport; WASMFS enabled for standalone builds (#36, #59) - Linux:
File::isOnRemovableDrive()implemented via/sys/block/<dev>/removable(#36)
zlibandoboeextracted from inline module sources into standalonethirdparty/modules for cleaner namespace isolation and build separation (#6, #7)TARGET_IDE_GROUPparameter onyup_standalone_appandyup_audio_plugin; all modules, tests, and examples placed in dedicated "Modules", "Tests", and "Examples" IDE folders (#11)appleFrameworks/appleWeakFrameworksmodule declaration fields unifying iOS and macOS framework lists (#36)- Per-platform C++ standard override via
*CppStandardmodule header fields (appleCppStandard,osxCppStandard,linuxCppStandard,wasmCppStandard,androidCppStandard,msftCppStandard) (#52) - Platform CMake files reorganized under
cmake/platforms/and loaded dynamically per target platform (#36) - Circular dependency detection for YUP modules (#111)
- Module link options support (per-platform
*LinkOptions) (#53) - Module target aliases (
yup::yup_core, etc.) (#53) - Code coverage:
YUP_ENABLE_COVERAGE, codecov integration (#54) - Test sharding support:
--gtest_total_shards/--gtest_shard_indexfor parallel CI runs (#119)
- Crash at startup when height/width is 0 on custom-scaled screens (#21)
- Application never quits: incorrect
quitMessagePostedordering instopDispatchLoop()(#42) - Redraw issues and app icon rendering on macOS (#31)
- iOS toolchain: removed hardcoded
DEVELOPER_DIRpath (#72) - CoreAudio thread safety: atomic operations replacing mutex (#76)
- SMPTE timecode validation, SSE macro, and memory fixes (#78)
- ZIP timestamp: missing
>>1for 2-second resolution (#92) StyledText::clear()fully resets state; caret bounds and glyph index for empty lines (#96)- Duplicated SDL symbols in Apple plugins (#112)
ComboBoxpopup re-opens on click-to-dismiss; addedignoreMouseDownAfterPopupDismissal(#114)AudioDeviceManagerdestructor race onmidiCallbackLock(#115)- Android oboe:
__ANDROID__preprocessor instead ofANDROID(#9) Graphics::drawImage(),renderStrokePath(),renderFillPath(),renderFittedText(): opacity not propagated to renderer - fixed (#119)SIMDRegister: tail-loop bounds check preventing out-of-bounds access inloadandstorepaths (#119)- Mouse-wheel events now dispatched to the component under the cursor when no component is focused (#30)
- Linux
Watchdog: inotify fd set to non-blocking, read buffer heap-allocated, thread join order corrected to prevent crash on destruction (#36)