Skip to content

Commit af24da2

Browse files
authored
Descend math, str.startswith/endswith and abs/hash/ord/min/max through gateways; retire their folds (#1964)
* math: descend math.sqrt through an interp2app gateway, drop the sqrt fold `math.sqrt` is now installed as `__majit_wrap_math_sqrt` (fixed arity 1), the interp_math.py `math1(space, math.sqrt, w_x)` / `_get_double` body with ll_math.py `ll_math_sqrt`'s branches: exact float or machine int, `x >= 0.0`, `isfinite(x)` -> `_float_sqrt(x)` (`sqrt_nonneg`, oopspec `math.sqrt_nonneg`); +inf returns the operand; every other shape goes to the `dont_look_inside` `sqrt_slow`, which runs the old body behind the declared-arity check. The generic builtin-call descent walks the gateway, so `try_walker_specialize_math_sqrt`, `try_walker_orthodox_float_sqrt`, `FLOAT_SQRT_DESCENT`, `MathFloatDomain::NonNegativeFinite` and `is_math_sqrt_function` go. Instructions retired, dynasm, upstream/main -> this tree: math_sqrt_hot.py 4,266,475,480 -> 4,189,283,971 math_folds_hot.py 5,248,358,508 -> 5,269,222,715 spec_folds! rows: 93 -> 92 fn try_walker_specialize_: 80 -> 79 Assisted-by: Claude * jit-trace: restore the binary_slice_str fold The synth-only census that retired it missed `pyre/extra_tests/parity_tests/binary_slice_str_specialization.py` and one other parity test, where it fires 6 times: `str[start:stop]` on an exact `str` with exact-int bounds. `binary_slice` still lowers to the `RuntimeHelperKind::BinarySlice` MayForce residual, so without the fold that loop keeps the opaque call. The other six folds retired with it fire in neither the synth nor the parity-test census. spec_folds! rows: 92 -> 93 fn try_walker_specialize_: 79 -> 80 Assisted-by: Claude * math: descend the math builtins through interp2app gateways, drop their folds Every math builtin except frexp is now an interp2app gateway (`__majit_wrap_math_*`) whose fast arm reads an exact float or machine int, pins the domain before the C call, and calls one unboxed leaf in descroperation; anything else runs the old body through a `dont_look_inside` `_slow`. The generic builtin descent walks the gateways, so the math_log_trig, math_float1, math_float2, math_fabs, math_isclose, math_ldexp, math_isqrt, math_floor, math_ceil and math_trunc folds and their helpers are removed. The fast arms pin overflow before the C call (exp/expm1 x < 709, exp2 x < 1023, sinh/cosh |x| < 709, pow via the frexp exponent bound) instead of testing the result, because the `_slow` fallback has no executable address once the call has run and the descent declines. ldexp takes only the arm where scaling is exact, and `_float_ldexp_raw` builds 2**exp from its bits instead of `2.0.powf(exp)`, which returned 0 for ldexp(1024.0, -1080). asinh/acosh/atanh call pymath (libm) through elidable raw leaves; std's formulas differ from libm in the last place. math_frexp keeps its fold: the gateway has to root the mantissa box across the exponent box, and `push_roots` does not lower in a walked body. frexp is installed from extra_init; `math_builtin_name` now answers only "frexp", by its BuiltinCode function pointer. The boxing leaves keep their arithmetic in their own bodies; the pow and ldexp gateways share the frexp-exponent and exact-ldexp bit arithmetic with them through local macros. wasm32 links no BUILTIN_WRAPPER_DESCRIPTORS, so publish_optional_fnaddrs binds every math gateway's descriptor path to its address there (math_gateway_fnaddrs), the way jit_fnaddr binds the interpreter's own gateways; without it the descent found no jitcode and import_math ran at 31x dynasm. Removed helpers with no remaining caller: trace_opcode's ccall_pow, float_pow_jit, sqrt_nonneg_jit and math_{log,cos,sin}_*_jit, walker_float_helper_addrs, and interp_math's jit_math_frexp_*, jit_math_ldexp_raw, jit_math_isclose_default and jit_math_isqrt_i64, and the math1_gamma_result_finite optional-module hook. spec_folds! rows 93 -> 83, try_walker_specialize_ fns 80 -> 72. Assisted-by: Claude * math: make the gateway leaves lift in the rtyper prepass The math gateways joined the rtyper skip-subject ratchet with 20 names: each wrapper failed Phase A because a leaf it calls did not lift. - `f64::abs` and `f64::to_int_unchecked::<i64>` reach the flowspace adapter as the unary ops `abs` and `cast_float_to_int`, which `normalize_unary_op_name` refused. `abs` is `operation.py`'s `abs` (`FloatRepr.rtype_abs`); `cast_float_to_int` is what `FloatRepr.rtype_int` emits for `int(x)`, so it maps to `int`. - The frexp/ldexp bit arithmetic uses `f64::to_bits` / `f64::from_bits`, which the front lowers to `longlong2float` calls, instead of `transmute`. - erf, erfc, ulp, gamma, lgamma, remainder and fmod call one `#[majit_macros::elidable]` raw leaf each (`ll_math.py` llexternals are `elidable_function=True`); float `%` and the pymath calls stay inside those leaves. Locally the darwin ratchet no longer lists any math name, and 40 previously skipped graphs (`_float_abs`, `_float_math1`, `_int_from_*`, `abs_structural`, ...) now lift. Instructions retired, dynasm, origin/main 4f8c388 -> this change: erf 0.687, erfc 0.689, gamma 0.715, lgamma 0.705, remainder 0.416, ldexp 0.92; fmod, ulp, frexp, pow, fabs, ceil, floor, trunc, isqrt, isclose within +-1.3%. Output is identical. spec_folds! rows 83 -> 83, try_walker_specialize_ fns 72 -> 72. Assisted-by: Claude * str: descend startswith/endswith through their gateways, drop their folds `__majit_wrap_str_descr_startswith` / `_endswith` called the arity and keyword checks before the fast arm. `reject_kwargs` reaches `w_dict_str_entries_wtf8`, which does not lower, so the builtin descent always declined and the `str_startswith` / `str_endswith` folds answered. The fast arm now runs first (two operands, both exact `str`, no tuple): a keyword dict rides the same slice and is never a `str`. It calls the elidable `unicodeobject::startswith` / `endswith` leaf instead of the inline `rstring_prefix_eq!` byte loop; walking that loop in a descended body gave wrong counts for a mixed hit/miss input. Every other shape runs `str_method_startswith` / `str_method_endswith` through a `dont_look_inside` slow path. Removed `try_walker_specialize_str_prefix_match` and its two residual-call gates. spec_folds! rows: 83 -> 81 fn try_walker_specialize_: 72 -> 71 Assisted-by: Claude * builtins: descend abs/hash/ord/min/max through their gateways, drop builtin_fold1/2 Each builtin is now installed as a `__majit_wrap_builtin_*` gateway whose fast arm pins an exact operand before any call and whose other shapes run the original body through a `dont_look_inside` slow path: - `abs`: an exact int other than `i64::MIN` calls `_int_abs`; an exact float calls `_float_abs`; an exact complex calls `complex_abs`, which keeps the `_float_math1` hub (and the `frexp` / `gamma` leaves it mints) reachable now that `builtin_abs` sits behind the slow path. - `hash`: an exact int computes `_hash_int` inline; an exact `str`, long, bytes or non-NaN float calls the `elidable_cannot_raise` leaf `hash_exact_scalar`. The `str` arm of `hash_value` moved into `str_hash_value` so the leaf does not reach `hash_value`'s tuple arm. - `ord`: the `elidable_cannot_raise` leaf `ord_exact_str_char` returns the code point of a one-code-point exact `str`, or -1. - `min` / `max`: two exact ints or two exact floats compare inline and return the winning operand. `is_builtin_hash_function` / `is_builtin_ord_function` compare against the new gateways. `hash`, `ord`, `min` and `max` publish their descriptors through `builtin_wrapper_descriptor!`. Removed `try_walker_specialize_builtin_fold1`, `try_walker_specialize_builtin_fold2`, their residual-call gates and helpers, and the `jit_builtin_folds` raw helpers except `is_repr_builtin` / `builtin_code_fn_of`. `builtin_folds_hot.py` drops its `spec-folds=` header. spec_folds! rows: 81 -> 79 fn try_walker_specialize_: 71 -> 69 Assisted-by: Claude * math: publish the gateway descriptors through builtin_wrapper_descriptor! Each `__majit_wrap_math_*` descriptor is now declared with `pyre_interpreter::builtin_wrapper_descriptor!`, which appends a linkme slice element natively and registers the same path through a constructor on wasm32. `math_gateway_fnaddrs` and the wasm32 loop in `publish_optional_fnaddrs` that bound those paths by hand are removed. spec_folds! rows: 79 -> 79 fn try_walker_specialize_: 69 -> 69 Assisted-by: Claude * jit-trace: drop the binary_slice_str fold again Reverts f2e7eb7. `str[start:stop]` on an exact `str` with exact-int bounds goes back to the `RuntimeHelperKind::BinarySlice` residual, as on main since #1939. design.md's census numbers follow. spec_folds! rows: 79 -> 78 fn try_walker_specialize_: 69 -> 68 Assisted-by: Claude * math: check frexp's arity in its body; frexp fold declines before guarding `frexp` is installed through the arity-1 `install` closure without the `py_checked_arity_fn!` wrapper the `py_module!` table gave it, so `math.frexp(1, 2)` and keyword calls reached the body. The body now runs `check_declared_positional_arity("frexp", 1, args)` itself, which keeps its function pointer the one `math_builtin_name` recognises. `try_walker_specialize_math_frexp` now runs the authoritative-executor, exact-type and finite/non-zero/normal checks (`frexp_fold_operand`) before recording the callable `GuardValue` and the operand coercion. `_float_ldexp_raw` documents that its caller has checked `-1022 <= exp <= 1023`. spec_folds! rows: 78 -> 78 fn try_walker_specialize_: 68 -> 68 Assisted-by: Claude * result_exc: skip the from_exc_object rebuild of a discarded Err `err_payload_is_dead` answers true when the Result shell reaches a block that neither reads nor forwards it (`let _ = f();`). `catch_and_rewrap` then keeps the caught word instead of calling `from_exc_object`, whose `PyObject` parameter does not union with the caught `Exception`. That union was the prepass phase-A failure of `pyre_interpreter::error::chain_context` (`let _ = _break_context_cycle(exc, active);`). New unit test: catch_and_rewrap_does_not_rebuild_a_discarded_err. spec_folds! rows 73 -> 73, fn try_walker_specialize_ 66 -> 66. Assisted-by: Claude * design.md: recount the fold census on top of #1972 73 spec_folds! rows (4 descent), 66 try_walker_specialize_ functions (62/1/3), 18 try_walker_orthodox_ functions, 559 synth fixtures, 50 distinct rows named by 54 spec-folds= headers. spec_folds! rows 73 -> 73, fn try_walker_specialize_ 66 -> 66. Assisted-by: Claude * majit: pay the darwin rtyper skip-subject baseline down 61 subjects no longer skipped on darwin at this branch (the math/abs gateway leaves and wrappers, chain_context, and graphs #1974 already fixed); no additions. The linux baseline is unchanged. spec_folds! rows 73 -> 73, fn try_walker_specialize_ 66 -> 66. Assisted-by: Claude * pyre-object: pass str/bytes leaf operands as object pointers `jit_str_find`, `jit_str_rfind`, `jit_str_contains`, `jit_str_compare`, `jit_bytes_contains` and `jit_bytes_contains_byte` took their `W_UnicodeObject` / `W_BytesObject` operands as `i64`, and their interpreter callers passed `obj as i64`. The codewriter lowers that cast to `cast_ptr_to_int`, so a traced call carried the receiver as an Int box the GC map does not track, the shape #1975 fixed for the bounded find/rfind/count leaves. These leaves now take `PyObjectRef`, and `contains_str`, `contains_bytes_like`, the descroperation caller and the specialize.rs caller pass the objects directly. spec_folds! rows 72 -> 72, fn try_walker_specialize_ 65 -> 65. Assisted-by: Claude * design.md: recount the fold census on top of #1976 72 spec_folds! rows (4 descent), 65 try_walker_specialize_ functions (61/1/3), 560 synth fixtures, 49 distinct rows named by 54 spec-folds= headers. spec_folds! rows 72 -> 72, fn try_walker_specialize_ 65 -> 65. Assisted-by: Claude
1 parent d5e932b commit af24da2

26 files changed

Lines changed: 1529 additions & 3641 deletions

File tree

‎majit/majit-metainterp/src/call_descr.rs‎

Lines changed: 1 addition & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -332,8 +332,7 @@ pub const CANNOT_RAISE_NO_HEAP_EFFECT_INFO: EffectInfo = EffectInfo {
332332
/// `extraeffect >= EF_FORCES_VIRTUAL_OR_VIRTUALIZABLE`;
333333
/// `EF_ELIDABLE_CANNOT_RAISE` is 0. Elidable with `can_collect=false` is
334334
/// therefore the analyzer output for an `@jit.elidable` leaf whose collect
335-
/// analyzer clears it, which is the shape of every helper in
336-
/// `jit_builtin_folds`.
335+
/// analyzer clears it.
337336
///
338337
/// `EffectInfo.__new__` additionally empties `_write_descrs_*` for every
339338
/// `EF_ELIDABLE_*`. Those sets are already empty here, so no other field

‎majit/majit-translate/src/front/result_exc.rs‎

Lines changed: 52 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2032,7 +2032,8 @@ fn rewire_one_call_site(
20322032

20332033
/// The `Err` payload of the match `exit` reaches is never read.
20342034
///
2035-
/// `Err(_)` still switches on the discriminant. The handler does not use
2035+
/// Either the shell is discarded unread (`let _ = f();`), or `Err(_)` still
2036+
/// switches on the discriminant. The handler does not use
20362037
/// the caught carrier, so rebuilding it would pass an `Exception` into
20372038
/// `from_exc_object`'s `PyObject` parameter.
20382039
fn err_payload_is_dead(graph: &FunctionGraph, exit: &Link, r: &Variable) -> bool {
@@ -2046,6 +2047,15 @@ fn err_payload_is_dead(graph: &FunctionGraph, exit: &Link, r: &Variable) -> bool
20462047
let Some(shell) = graph.blocks[exit.target.0].inputargs.get(pos).cloned() else {
20472048
return false;
20482049
};
2050+
// `let _ = f()?`-less discard: the shell reaches its block and nothing
2051+
// reads or forwards it, so neither arm's payload is ever observed.
2052+
let uses = count_var_uses(graph, &shell);
2053+
if uses.op_uses == 0
2054+
&& uses.link_uses == 0
2055+
&& !matches!(&graph.blocks[exit.target.0].exitswitch, Some(ExitSwitch::Value(v)) if *v == shell)
2056+
{
2057+
return true;
2058+
}
20492059
let Ok((_, _, disc_shell)) = match_discriminant(graph, exit.target.0) else {
20502060
return false;
20512061
};
@@ -5267,4 +5277,45 @@ mod rebuilt_shell_collapse_tests {
52675277
"a non-match consumer keeps the rebuilt shells"
52685278
);
52695279
}
5280+
5281+
/// `let _ = f();` discards the `Result` unread, so the caught word is
5282+
/// never materialised back into a carrier: `from_exc_object` takes a
5283+
/// `PyObject` and the caught word is an `Exception`.
5284+
#[test]
5285+
fn catch_and_rewrap_does_not_rebuild_a_discarded_err() {
5286+
let mut graph = FunctionGraph::new("rewrap_discarded");
5287+
let a = graph.startblock;
5288+
let r = graph
5289+
.push_op_var(
5290+
a,
5291+
OpKind::Call {
5292+
target: CallTarget::function_path(["callee"]),
5293+
args: Vec::new(),
5294+
result_ty: ValueType::Ref(None),
5295+
},
5296+
true,
5297+
)
5298+
.expect("call");
5299+
let (tail, _) = graph.create_block_with_arg_vars(1);
5300+
graph.set_return(tail, None);
5301+
graph.set_goto(a, tail, vec![r.clone()]);
5302+
let spec = crate::ErrorCarrierSpec {
5303+
carrier_path: "carrier::PyError",
5304+
carrier_wrappers: &[],
5305+
to_exc_object: None,
5306+
from_exc_object: Some(("PyError", "from_exc_object")),
5307+
};
5308+
catch_and_rewrap(&mut graph, a.0, &r, "<(),PyError>", &ValueType::Void, spec)
5309+
.expect("rewrap");
5310+
let rebuilds = graph
5311+
.blocks
5312+
.iter()
5313+
.flat_map(|b| &b.operations)
5314+
.filter(|op| {
5315+
matches!(&op.kind, OpKind::Call { target, .. }
5316+
if format!("{target:?}").contains("from_exc_object"))
5317+
})
5318+
.count();
5319+
assert_eq!(rebuilds, 0, "a discarded Err payload is not rebuilt");
5320+
}
52705321
}

‎majit/majit-translate/src/translator/rtyper/flowspace_adapter.rs‎

Lines changed: 13 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -589,11 +589,22 @@ fn normalize_unary_op_name(source_name: &str) -> Result<String, TyperError> {
589589
// string-build lowering (`str(arg)` ++ `ll_strconcat`) in place of
590590
// the graph-less `fmt::rt::Argument::new_display` chain.
591591
"str" => Ok("str".to_string()),
592+
// `abs` — `operation.py add_operator('abs', 1, dispatch=1,
593+
// pyfunc=abs, pure=True)`, dispatched at `rtyper.rs "abs"` into
594+
// `FloatRepr.rtype_abs` (`float_abs`) / `IntegerRepr.rtype_abs`.
595+
// The front lowers `f64::abs` to this unary op.
596+
"abs" => Ok("abs".to_string()),
597+
// `cast_float_to_int` is what `rfloat.py FloatRepr.rtype_int`
598+
// emits for `int(x)` on a float, so the flowspace op it stands for
599+
// is `int` (`operation.py add_operator('int', 1, dispatch=1,
600+
// pyfunc=int)`). The front emits it for
601+
// `f64::to_int_unchecked::<i64>` and the float-to-int scalar cast.
602+
"cast_float_to_int" => Ok("int".to_string()),
592603
other => Err(TyperError::missing_rtype_operation(format!(
593604
"normalize_unary_op_name: pyre UnaryOp `{other}` has no \
594605
flowspace counterpart (operation.py registers \
595-
`pos` / `neg` / `invert` / `bool` and the ported `str` \
596-
as unary ops; \
606+
`pos` / `neg` / `invert` / `bool` / `abs` and the ported \
607+
`str` as unary ops, and `cast_float_to_int` stands for `int`; \
597608
`same_as` is rtyper's internal renaming op per \
598609
rtyper.py:478-481; all 13 typed cast names retired \
599610
across Slices A.3 / B.1 / A.4a / A.4b / A.4c — frontend \

‎majit/rtyper-skip-subjects.darwin.txt‎

Lines changed: 0 additions & 32 deletions
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,6 @@
44
# done-when.
55
# corpus=c63631cace2c58a8 platform=darwin
66
two-phase-never-a-subject eval_loop_jit_portal
7-
two-phase-never-a-subject ll_int_abs
87
two-phase-never-a-subject ll_listslice_minusone__ref__majit__object_ref_gcarray
98
two-phase-never-a-subject ll_listslice_startonly__ref
109
two-phase-never-a-subject ll_listslice_startonly__ref__majit__object_ref_gcarray
@@ -150,21 +149,17 @@ two-phase-never-a-subject pyre_interpreter::baseobjspace::view_as_kwargs
150149
two-phase-never-a-subject pyre_interpreter::baseobjspace::w_type_getdictvalue
151150
two-phase-never-a-subject pyre_interpreter::builtins::<Impl>::acquire
152151
two-phase-never-a-subject pyre_interpreter::builtins::<Impl>::as_mut_slice
153-
two-phase-never-a-subject pyre_interpreter::builtins::__majit_wrap_builtin_abs
154152
two-phase-never-a-subject pyre_interpreter::builtins::__majit_wrap_builtin_dunder_import
155153
two-phase-never-a-subject pyre_interpreter::builtins::__majit_wrap_builtin_isinstance
156154
two-phase-never-a-subject pyre_interpreter::builtins::__majit_wrap_builtin_len
157155
two-phase-never-a-subject pyre_interpreter::builtins::__majit_wrap_builtin_locals
158-
two-phase-never-a-subject pyre_interpreter::builtins::abs_structural
159156
two-phase-never-a-subject pyre_interpreter::builtins::acquire_readbuf
160157
two-phase-never-a-subject pyre_interpreter::builtins::backing_exports_incref
161158
two-phase-never-a-subject pyre_interpreter::builtins::bare_super_frame_layout
162159
two-phase-never-a-subject pyre_interpreter::builtins::bare_super_frame_layout_words
163160
two-phase-never-a-subject pyre_interpreter::builtins::base_exception_str_method
164161
two-phase-never-a-subject pyre_interpreter::builtins::bind_pos_or_kw
165162
two-phase-never-a-subject pyre_interpreter::builtins::buffer_export_incref
166-
two-phase-never-a-subject pyre_interpreter::builtins::builtin_abs
167-
two-phase-never-a-subject pyre_interpreter::builtins::builtin_abs_obj
168163
two-phase-never-a-subject pyre_interpreter::builtins::builtin_float
169164
two-phase-never-a-subject pyre_interpreter::builtins::builtin_isinstance
170165
two-phase-never-a-subject pyre_interpreter::builtins::builtin_len
@@ -257,15 +252,13 @@ two-phase-never-a-subject pyre_interpreter::error::<Impl>::normalize_exception
257252
two-phase-never-a-subject pyre_interpreter::error::<Impl>::os_error_syscall
258253
two-phase-never-a-subject pyre_interpreter::error::<Impl>::os_error_syscall2
259254
two-phase-never-a-subject pyre_interpreter::error::<Impl>::os_error_with_errno
260-
two-phase-never-a-subject pyre_interpreter::error::<Impl>::record_context
261255
two-phase-never-a-subject pyre_interpreter::error::<Impl>::render_exception_wtf8
262256
two-phase-never-a-subject pyre_interpreter::error::<Impl>::set_cause
263257
two-phase-never-a-subject pyre_interpreter::error::<Impl>::to_exc_object
264258
two-phase-never-a-subject pyre_interpreter::error::<Impl>::write_to_sys_stderr
265259
two-phase-never-a-subject pyre_interpreter::error::<Impl>::write_unraisable
266260
two-phase-never-a-subject pyre_interpreter::error::<Impl>::write_unraisable_default
267261
two-phase-never-a-subject pyre_interpreter::error::<Impl>::write_unraisable_with_traceback
268-
two-phase-never-a-subject pyre_interpreter::error::chain_context
269262
two-phase-never-a-subject pyre_interpreter::error::emit_report_to_host_stderr
270263
two-phase-never-a-subject pyre_interpreter::error::exc_object_class_name
271264
two-phase-never-a-subject pyre_interpreter::error::system_error_from_cause
@@ -940,24 +933,6 @@ two-phase-never-a-subject pyre_interpreter::module::thread::rlock_class::__majit
940933
two-phase-never-a-subject pyre_interpreter::module::thread::rlock_class::__majit_wrap_acquire_timed
941934
two-phase-never-a-subject pyre_interpreter::module::thread::rlock_class::__majit_wrap_release
942935
two-phase-never-a-subject pyre_interpreter::module::time::interp_time::duration_since_epoch
943-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_abs
944-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_erf
945-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_erfc
946-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_fmod
947-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_frexp_mantissa
948-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_gamma
949-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_isclose
950-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_lgamma
951-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_math1
952-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_math2
953-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_remainder
954-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_float_ulp
955-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_int_frexp_exponent
956-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_int_from_ceil
957-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_int_from_float
958-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_int_from_floor
959-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_int_from_trunc
960-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::_int_isqrt
961936
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::add
962937
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::add_impl
963938
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::and_
@@ -966,10 +941,7 @@ two-phase-never-a-subject pyre_interpreter::objspace::descroperation::bytes_conc
966941
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::bytes_concat_type_error
967942
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::compare
968943
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::compare_slot
969-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::complex_abs
970944
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::complex_pow
971-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::complex_quot
972-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::complex_truediv
973945
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::concat_operand_name
974946
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::float_pow_impl
975947
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::float_pow_inner
@@ -999,7 +971,6 @@ two-phase-never-a-subject pyre_interpreter::objspace::descroperation::pos
999971
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::pos_inner
1000972
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::pow
1001973
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::pow_binary
1002-
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::real_complex_quot
1003974
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::rshift
1004975
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::rshift_impl
1005976
two-phase-never-a-subject pyre_interpreter::objspace::descroperation::sequence_repeat
@@ -1246,9 +1217,6 @@ two-phase-never-a-subject pyre_interpreter::shared_opcode::opcode_store_subscr
12461217
two-phase-never-a-subject pyre_interpreter::type_methods::__majit_wrap_dict_descr_items
12471218
two-phase-never-a-subject pyre_interpreter::type_methods::__majit_wrap_dict_descr_keys
12481219
two-phase-never-a-subject pyre_interpreter::type_methods::__majit_wrap_dict_descr_values
1249-
two-phase-never-a-subject pyre_interpreter::type_methods::__majit_wrap_str_descr_endswith
1250-
two-phase-never-a-subject pyre_interpreter::type_methods::__majit_wrap_str_descr_startswith
1251-
two-phase-never-a-subject pyre_interpreter::type_methods::arity_at_least
12521220
two-phase-never-a-subject pyre_interpreter::type_methods::arity_at_most
12531221
two-phase-never-a-subject pyre_interpreter::type_methods::arity_exact
12541222
two-phase-never-a-subject pyre_interpreter::type_methods::builtin_encoder

‎pyre/bench/synth/builtin_folds_hot.py‎

Lines changed: 26 additions & 62 deletions
Original file line numberDiff line numberDiff line change
@@ -2,70 +2,34 @@
22
# Ubuntu run 33279264115: 1.7-2.2x; the ceiling is twice the slowest,
33
# rounded up to one decimal place.
44
# pyre-check: skip-cpython
5-
# Fitted between the two arms. With every fold in place the three runners read
6-
# 4.6x / 4.7x on darwin-arm64, 7.2x / 7.6x on ubuntu-24.04 and 9.2x / 10.0x on
7-
# windows, at half the loop counts below. The windows pair was marked `?`
8-
# there: pypy's execution-only time sat under FLOOR_GATE_MIN_BASELINE_S, which
9-
# is 0.15625s on a host whose CPU accounting advances in 1/64s ticks, and the
10-
# pair cleared the ceiling of 8 it carried then only because `_compare_buffer`
11-
# grants two ticks per unit of limit. Doubling the work carries that baseline
12-
# over the bar, so the number is judged rather than excused. With
13-
# PYRE_FBW_NO_SPECIALIZE=builtin_fold1,builtin_fold2 putting the same loops back
14-
# on the residual, darwin-arm64 reads 52.9x and 61.4x -- eleven times the folded
15-
# arm on that host. The gate clears the highest folded reading by 20% and still
16-
# sits an order of magnitude under the residual arm scaled to it.
5+
# The ceiling was fitted when the `builtin_fold1` / `builtin_fold2` hand
6+
# folds answered these calls, at 4.6x / 4.7x on darwin-arm64, 7.2x / 7.6x on
7+
# ubuntu-24.04 and 9.2x / 10.0x on windows (half the loop counts below). The
8+
# loop counts are doubled so pypy's execution-only time clears
9+
# FLOOR_GATE_MIN_BASELINE_S on the windows runner, whose CPU accounting
10+
# advances in 1/64s ticks, and so the fixed startup-subtraction error is a
11+
# smaller share of the denominator.
1712
#
18-
# Those three-runner readings predate two changes to how these folds are
19-
# recorded. The `Int1` and `Float1` channels became elidable calls, so a fold
20-
# whose operand does not change is hoisted out of the loop instead of called
21-
# once per iteration. And `min` / `max` over a pair of exact machine ints
22-
# stopped emitting a call at all: what is guarded is the ordering of the two
23-
# unboxed values, and under that ordering the answer is the winning operand's
24-
# own reference. darwin-arm64 reads 2.0x / 2.2x where it read 4.6x / 4.7x.
13+
# Every builtin here is now an interp2app gateway the generic builtin descent
14+
# walks, the way the tracer walks straight into an RPython builtin body:
2515
#
26-
# pyre-check: spec-folds=builtin_fold1,builtin_fold2
27-
# This fixture carried a wasm allowance of 13 while every folded builtin still
28-
# left the trace module: a fold removes the frame force, the argument rooting,
29-
# the execution-context resolution and the gateway binding, but each one still
30-
# lowered to a call into a raw helper, and `abs(x)`'s helper is `(i64) -> f64`,
31-
# a mixed signature the backend would not lower without a caller vouching for
32-
# it. 31,998,958 crossings, 100% of them that one helper, 3.3s of a 5.0s run.
33-
# Vouching it (`float_fold_helper_addrs`) takes the crossings to zero and the
34-
# fixture to 1.2x, so the allowance is gone rather than lowered.
16+
# hash_int `_hash_int` inline on an exact machine int.
17+
# hash_str an exact `str` pinned first, then one
18+
# `elidable_cannot_raise` leaf the optimizer hoists out
19+
# of the loop.
20+
# ord_str the same shape, the leaf returning the code point of a
21+
# one-code-point exact `str`.
22+
# abs_int/abs_float `descr_abs` on an exact int (short of `i64::MIN`) or
23+
# float, boxed by a leaf the optimizer keeps virtual.
24+
# min_max two exact machine ints or floats compared inline; the
25+
# answer is the winning operand's own reference.
3526
#
36-
# The loop counts are twice what they once were, and stay that way: the
37-
# denominator is small enough for startup subtraction to move it, and the
38-
# subtraction error is a fixed number of milliseconds, so doubling the work
39-
# halves its share. Every recorded jit-stats counter is unchanged by it.
40-
#
41-
# `spec-folds` is the exact instrument neither ratio is. Six loops sum into
42-
# one number, so retiring one channel moves it by less than the span this
43-
# fixture reads across the runners; the census gates each fold's coverage
44-
# instead, and it reads the same on every host.
45-
#
46-
# Every builtin the generic walker fold covers, one hot loop per channel.
47-
#
48-
# Without a fold, `hash(x)` / `ord(c)` / `abs(x)` / `min(a, b)` each reach the
49-
# interpreter as `bh_call_fn(builtin, NULL, ...)`, and that residual costs the
50-
# same for all of them: the frame force, the argument rooting, the execution
51-
# context resolution and the gateway signature binding all run before the body
52-
# does. The fold emits a direct call into the builtin's raw helper, a guard on
53-
# that channel's decline sentinel, and an inline `wrapint` / `wrapfloat` the
54-
# optimizer can keep virtual.
55-
#
56-
# The ratio is the detector here: losing a fold changes no jit-stats counter,
57-
# because the residual it falls back to compiles the same loop.
58-
#
59-
# hash_int/hash_str the `Int1` channel — an `i64` result and the
60-
# `INT_FOLD_DECLINE` guard.
61-
# ord_str the same channel on an operand whose acceptance is a
62-
# length, not a type.
63-
# abs_int/abs_float one builtin holding two rows, one per result channel.
64-
# min_max the `Ref2` channel — the helper returns one of its own
65-
# arguments, so nothing is allocated at all. A pair of
66-
# exact machine ints does not reach the helper: the
67-
# ordering guard picks the winner and the answer is that
68-
# operand's own reference.
27+
# Every other operand shape reaches the original builtin body through a
28+
# `dont_look_inside` slow path. Without the descent each call is a
29+
# `bh_call_fn(builtin, NULL, ...)` residual: the frame force, the argument
30+
# rooting, the execution-context resolution and the gateway signature binding
31+
# all run before the body does, and the ratio is the detector, because the
32+
# residual compiles the same loop and changes no jit-stats counter.
6933
HASH_N = 32000000
7034
ORD_N = 32000000
7135
ABS_N = 32000000
@@ -84,7 +48,7 @@ def run_hash_int():
8448
def run_hash_str():
8549
# A string's hash is seeded per process, so the digest itself cannot be
8650
# printed. Count the iterations that agree with the first one instead:
87-
# the fold still has to produce the digest, and the count is invariant.
51+
# the descent still has to produce the digest, and the count is invariant.
8852
s = "specialize"
8953
first = hash(s)
9054
same = 0

0 commit comments

Comments
 (0)