Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions pyre/pyre-interpreter/src/module/_ast/convert.rs
Original file line number Diff line number Diff line change
Expand Up @@ -4473,10 +4473,10 @@ fn push_literal(parts: &mut Vec<JoinedPart>, (start, end): (u32, u32), value: &W
/// `ast::ConstantValue::Str` and every identifier field are `str`, so a lone
/// surrogate has nowhere to go on the way back into the compiler. It gets
/// there through an ordinary `ast.parse` round trip, since the tree the parse
/// answers does carry one, and [`w_str_get_value`] would take the process
/// down over it.
/// answers does carry one; [`w_str_get_value_opt`] returns `None` and this
/// path raises `UnicodeEncodeError`.
///
/// [`w_str_get_value`]: pyre_object::w_str_get_value
/// [`w_str_get_value_opt`]: pyre_object::w_str_get_value_opt
fn utf8_only(value: PyObjectRef) -> AstResult<&'static str> {
if let Some(text) = unsafe { pyre_object::w_str_get_value_opt(value) } {
return Ok(text);
Expand Down
4 changes: 2 additions & 2 deletions pyre/pyre-object/src/dictmultiobject.rs
Original file line number Diff line number Diff line change
Expand Up @@ -5746,8 +5746,8 @@ pub unsafe fn w_module_dict_values_inner(obj: PyObjectRef) -> Vec<PyObjectRef> {
/// Keys that carry a lone surrogate (not valid UTF-8) are skipped:
/// the remaining `&str`-keyed consumers (dict_storage_store, module
/// `__dir__`, builtins-module iteration) cannot yet represent a
/// surrogate key, so skipping them here avoids the [`w_str_get_value`]
/// panic. The keyword-argument ABI no longer uses this helper — it
/// surrogate key, so skipping them here avoids a UTF-8 view of WTF-8
/// storage. The keyword-argument ABI no longer uses this helper — it
/// threads the byte-ish key through [`w_dict_str_entries_wtf8`].
/// # Safety
/// The caller must uphold every validity, runtime-type, aliasing, and lifetime
Expand Down
19 changes: 6 additions & 13 deletions pyre/pyre-object/src/unicodeobject.rs
Original file line number Diff line number Diff line change
Expand Up @@ -927,15 +927,6 @@ pub fn box_str_constant(value: &Wtf8) -> PyObjectRef {
obj
}

/// There is no panicking `&str` view of a `W_UnicodeObject`.
///
/// `W_UnicodeObject.text_w` returns `self._utf8` verbatim, surrogates
/// included, so a caller that only forwards the buffer uses
/// [`w_str_get_wtf8`]. A caller that needs valid UTF-8 uses
/// [`w_str_get_value_opt`] and reports `UnicodeEncodeError`
/// ("surrogates not allowed"), the strict utf-8 encoder's reason
/// (`unicodehelper.py`).

/// The `&str` view of a WTF-8 buffer already known to hold no lone
/// surrogate.
///
Expand Down Expand Up @@ -991,9 +982,9 @@ pub unsafe fn w_str_is_utf8(obj: PyObjectRef) -> bool {

/// Borrow the WTF-8 view of a known W_UnicodeObject, surrogate-aware.
///
/// Unlike [`w_str_get_value`], this never panics on lone surrogates.
/// Callers that must handle surrogate-bearing strings (codec encode,
/// repr) read code points through this accessor.
/// `W_UnicodeObject.text_w` returns `self._utf8` verbatim, surrogates
/// included. Callers that must handle surrogate-bearing strings (codec
/// encode, repr) read code points through this accessor.
///
/// # Safety
/// `obj` must point to a valid `W_UnicodeObject`.
Expand Down Expand Up @@ -1135,7 +1126,9 @@ pub unsafe fn w_str_hash_memoized(obj: PyObjectRef) -> i64 {
///
/// String-keyed fast paths that store keys in a `&str`-keyed map use this
/// to skip surrogate keys and fall through to the generic object-keyed
/// path instead of panicking in [`w_str_get_value`].
/// path. Callers that need valid UTF-8 report `UnicodeEncodeError`
/// ("surrogates not allowed"), the strict utf-8 encoder's reason
/// (`unicodehelper.py`).
///
/// # Safety
/// `obj` must point to a valid `W_UnicodeObject`.
Expand Down
Loading