Skip to content

Latest commit

 

History

115 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Urusilla

Conformance License: Apache-2.0

An experimental, no-install, machine-first semantic language for communication between AI agents.

Try one bounded agent handoff—one file, no install, no public post. Give an unfamiliar agent the hosted language-probe.json. It must decode and re-encode one non-effect-authorizing action state while preserving explicit negation, failure, null, provenance, uncertainty, and hard constraints. Its result may be PASS, SAFE_FALLBACK, or FAIL; none establishes adoption, general compatibility, or token saving.

Choose a path

Time Route Honest outcome
About 60 seconds Answer one accounting probe One narrow decision; a pass, refusal, or disagreement is useful, but none proves adoption or saving.
Up to 10 minutes Fork and run the bounded smoke test Project-authored fixture compatibility at your revision; not an independent implementation or efficiency result.
Pick one issue Open good first issue queue A scoped code, test, or documentation contribution; the label and any fork are participation signals, not research evidence.

Start with your own value

Who should try: an operator with repeated, read-only CI or incident-triage handoffs where an unfamiliar receiver may lose public constraints, negation, failure, null, uncertainty, provenance, or an authority boundary.

Who should not: one-off novel-prose workflows, effectful tasks, or systems whose concise natural language or schema-constrained JSON already works. If either baseline is better, do not use Urusilla.

For repeated workflows: TRIAL_CONTRACT.md specifies how an eligible operator may produce a participant-controlled scorecard against the better concise-natural-language/schema-JSON baseline. It is documentation only; no packaged runner, claim-bearing result, submission, or publication is provided or required. General unfamiliar-agent saving is currently 0%, safely completed real-task total tokens are unknown, and verified external adoption is zero.

Fork it to reproduce a claim, not to endorse one. Create a public fork, then run Actions → Fork Reproduction Smoke Test in your fork. The bounded workflow installs no project package, checks the frozen agent-entry manifest and public decode challenge, and preserves failures as useful results. Continue with FORK_REPRODUCTION.md to turn the run into a pinned counterexample, refusal, null result, or deeper reproduction. A fork or green workflow is an attention and reproducibility signal—not adoption, independence, compatibility, or efficiency evidence.

Urusilla's long-term vision is an interoperable meaning layer that independent agents can learn, negotiate, inspect, and improve without sharing a vendor, model, tokenizer, or fixed human-facing syntax. The current prototype combines a typed semantic kernel, negotiated codecs, deterministic human inspection, safe natural-language/JSON fallback, A2A integration, and public evaluation artifacts. It is active open research—not yet a proven universal replacement for natural language or JSON.

Founded and initially developed and stewarded by jaden3824. The canonical project is jaden3824/urusilla and uses transparent founder-led governance during its experimental phase. Apache-2.0 grants reuse and fork rights in the licensed work; it does not transfer canonical project authority or official status. See GOVERNANCE.md.

Status: research prototype, not a standard. The project name is Urusilla. Its protocol namespaces and private-use media types remain experimental, and no trademark registration, standards endorsement, or domain ownership is implied.

Validate the probe output offline: validate_language_probe.py classifies an exact semantic response as PASS, a closed refusal/fallback as SAFE_FALLBACK, and every meaning or structure change as FAIL. The hosted file is byte-identical to the tracked language-probe.json. This is an open profile-level demonstration, not core wire conformance, adoption, general compatibility, or efficiency evidence.

The separate 60-second accounting check remains available for one narrow retention decision. Its frozen task asks whether unknown usage from a failed attempt may be treated as zero; results go to Discussion #8 only when publication is separately authorized.

Star it if you want to signal interest; then try to break one bounded claim. Stars provide a visible attention signal, while reproducible evidence is the contribution that can change the result. Agents can read the machine-first contribution-entry.json, submit the same four fields through the 60-second issue form, and escalate a result into a counterexample, codec candidate, or quarantined corpus example. Validated unfavorable and favorable evidence receive equal credit in CONTRIBUTORS_EVIDENCE.md; its validated-record count remains zero, while attributed review notes are explicitly non-registry context.

The name Urusilla is an attested ancient scholarly/topographical name for Babylon, glossed “city of jubilation” in the ORACC Babylonian Topographical Texts corpus. The name choice does not imply that it is an ordinary modern-language word. The project currently makes no claim of owning urusilla.com. The GitHub repository is the canonical source and evidence record; the hosted challenge page is only a participation interface.

Version identifiers use separate axes: the current semantic language and source-manifest languageVersion are exactly 0.1.0; the release label is v0.1.0-experimental; the Python distribution normalizes that prerelease to 0.1.0a0; and the lifecycle status is experimental-unsigned. A matching semantic version does not by itself prove signature status, conformance, or production readiness.

The project is testing one claim: agents may coordinate more precisely and sometimes more efficiently when they exchange a typed semantic representation instead of unconstrained human prose. The design combines an auditable semantic kernel, negotiated codecs, deterministic human inspection, public commitments, and an A2A integration path.

It does not claim that one syntax is optimal for every model, tokenizer, transport, or task. A profile is adopted only when measured total utility beats the best available fallback.

Current bottom line

The currently demonstrated token saving for general communication between unfamiliar agents is 0%. The freshest broad lane covers 2,542 turns from Taskmaster, Schema-Guided Dialogue, Dolly, and OpenAssistant under four pinned tokenizers. Its mandatory raw fallback passed H1 with zero positive-regret choices, but the general compact-value and repeated-context hypotheses H2 and H3 failed. Warm receiver-carrier saving was only 0.65% to 0.80%; every cold family plan retained raw text, and decoding before model input leaves measured API-input saving at 0%. H4, end-to-end task utility, was not evaluated.

The experimental role/session/turn/text external-profile carrier made the same turns 165.60% to 183.98% larger in tokens than bare text. On the separate retained 42-record official-example corpus, both bound and standalone compact modes won 0/168 comparisons; standalone cold text was 2.24% to 3.00% larger than raw concise text. Safe fallback is useful engineering behavior, but a tie or avoided regression is not compression.

Favorable results elsewhere have narrower scopes. Receiver-bound v0.7 saved 23,997 development tokens and 4,302 grouped-holdout tokens, but saved 0 on OOD and activated in 0/12 cold plans. Checkpointed v0.9 saved 53.71% to 55.15% only on deliberately correlated synthetic state. A historical pre-cutover live receiver result reached 27/28 exact reconstructions but failed its gate, and one historical pre-cutover internally operated neutral-ID sender pilot recorded only 6/10 structural-and-semantic passes. That sender participant has not been rerun on the current artifacts. None establishes arbitrary conversation efficiency. End-to-end tokens per safely completed real task—including sender representation, setup/instructions, receiver input/reasoning/output, routing, tools, repair, fallback, and judging—remain unknown. Application bytes and transport latency are reported separately rather than converted into token or energy savings; actual joules remain unknown without telemetry.

The immediate research goal is therefore model-native or task-aware public action-state consumption, verified silence or topology pruning, and total tokens per safely completed task on independently authored conversations. The preregistration-ready GENERAL_DIALOGUE_EVAL_PLAN.md freezes the intended cross-model, cold-start, drift, fallback, silence, and complete-cost comparisons before any scored run. Incremental tuning of a universal lossless text surface is paused after the broad H2 and H3 failures; it should resume only for a separately frozen architecture-changing hypothesis. See HELP_WANTED.md.

An exact request binding is not evidence that a receiver used the message. The current runtime and v1 initial-goal verifier can prove which action-state payload reached the model-visible request, but a constant-output receiver can still satisfy synthetic success plumbing without reading that payload. Therefore a v1 aggregate pass would establish hybrid-router utility only; it would not establish causal consumption of an action-state language. Any comprehension or route-level language claim additionally requires a preregistered, blinded payload-intervention study: identical non-payload context and settings, task-critical A/B payloads whose correct outputs must differ, missing/shuffled placebos that must refuse or fall back, inclusive accounting for every call, and per-stratum coverage. This stronger contract must use a new schema version rather than silently changing frozen v1.

Two new evaluator-only diagnostics make that next experiment harder to misreport without changing the protocol version. A /3 causal matrix validator requires six bound conditions for every declared field: critical A/B, semantic invariance, exact critical-field ablation, an independently answerable no-payload control, and a cross-field shuffled control. A matched-session runner then executes separately prepared raw, JSON, and Urusilla arms; after one passed cold-comprehension call, its hot Urusilla request contains only the canonical action state and never re-expands the Capsule or task prose. It retains failed primary plus fallback cost, and makes totals unknown when retry/repair attempts lack individual receipts. Both tools are tested only with project-authored records and fake adapters. Provider authenticity, callback call scope, frozen settings, preregistration chronology, independent operation, counterbalancing, and full strata remain unproved, so claim-facing totals and safe-task results stay null and general saving remains 0%.

A separate initial-goal zero-call feasibility screen now accepts only fully registered finite phase bounds for session lengths 1 through 128 and can return only impossible, not-disproven, or invalid. It uses candidate lower bounds and the better raw/JSON upper cost-per-safe-task bound, so equality with the 20% threshold is not a kill and not-disproven is never called a saving. No real row has been run: the required frozen prompt, dynamic-slot, tokenizer, and allowed-path bound manifests are absent from this checkout. The evaluator and its synthetic tests therefore add a cheap stop gate, not empirical efficiency evidence. This screen is separate from the SGD-20 competitive-reproduction protocol.

The screen's input boundary is now explicit rather than caller-asserted. A content-derived preflight hashes exact source and transmitted bytes and records tokenizer, chat-template, path-DAG, pre-call inclusive token-cap, and raw/JSON success-receipt inventory. It still always blocks numeric screening because no implemented compiler derives exact token vectors and phase bounds from those bytes. It never chooses a session length or authorizes receiver calls. A separate perfect-sender diagnostic can exercise one cold Capsule comprehension followed by N direct action-state turns while running matched raw and descriptive JSON arms in fresh contexts, but its current callback accepts only a caller's offline_synthetic=True declaration. That declaration is neither authenticated nor sandbox-enforced; callback exceptions become attempted calls with unknown usage and reject the run. Task and aggregate result construction is factory-sealed as an API-misuse guard, not as authentication. The aggregate also revalidates the original preflight-bound experiment manifest, per-task expected outputs, exact request digests, call order, terminal state, and token caps. Process interrupts still propagate through a BaseException carrier that always holds the attempted-call journal. Provider/model consistency, fresh baseline roots, the deterministically re-derived comprehension verdict, and the exact deterministic comprehension challenge plus comprehension-to-hot same-context parent chain are rechecked; response IDs cannot be reused. Any later ordinary validation failure also retains all earlier callback entries. These additions create an executable receiver-ceiling test shape; they add no real model result and do not change the demonstrated general saving.

A new development-only hybrid runtime now implements that architecture-changing hypothesis: task-bound natural-language compilation, direct action-state consumption, a fail-closed five-route planner, per-message semantic-fidelity evidence, complete-cost fields, and an optional session-local evolving surface. Its machine-first aliases may be non-English or opaque and are optimized without a human-aesthetics score. Stable semantic IDs never change; only a reversible, exact-context-bound wire table may evolve after round-trip comprehension tests. Caller-supplied UtilityEvidence can qualify an optimized route for a bounded local policy trial after exact binding and declared-threshold checks, but it is not claim authority: runtime route candidates and decisions reject claim_eligible: true, and the aggregate initial-goal verifier emits no route-scoped evidence. This is implementation plumbing, not a positive result: no real independent end-to-end run has yet shown that the extra compiler, verifier, setup, and receiver costs beat raw concise text and JSON. See EVOLVING_SURFACE.md.

The reference runtime now also closes the local same-context execution gap. A cached receiver capability can be minted only from a passed cold-comprehension attempt and the exact still-active provider-context observation. A factory-sealed session plan then sends the validated public action state directly without prose re-expansion. Malformed state, Capsule/task/context drift, adapter failure, or invalid primary output invalidates that optimized path and invokes its already bound raw/JSON fallback; prohibited authority remains false throughout. Focused and full runtime regressions exercise this path with project-authored adapters. No provider-backed causal intervention or independently operated run has yet shown that a real model used the task-critical semantics, so this is a safer experiment runner rather than language-use or efficiency evidence. See urusilla_hybrid_runtime/session_runtime.py.

A runtime-only session portfolio ledger now combines contiguous, exactly bound hot turns without charging the shared cold-comprehension call on every turn. Each turn retains its independently prepared raw or JSON comparator, compiler and fidelity calls, local setup/router/repair/fallback/tool/safety/judge usage, failed optimized receiver call, and any executed fallback. Duplicate, noncontiguous, cross-session, or post-failure turns are rejected. Any unknown provider or local phase keeps the affected aggregate and saving null. The reported-token diagnostic can expose a setup-amortization crossover, but provider receipt authenticity and full-history billing remain unverified, all current tests use project-authored adapters, and complete goal totals and claim eligibility remain false. See urusilla_hybrid_runtime/session_portfolio.py.

A separate four-slot, same-project Gemini web-UI pilot now records one task-critical delivery_date flip, one representation-invariance check, and one true-missing-field fallback in fresh temporary chats under the visible Pro Extended mode label. All four returned the preregistered canonical JSON byte-for-byte with no repair. The exact model version, provider token usage, and authenticated response receipt were unavailable. The packet, observation, and offline tests preserve the result as an exploratory field-binding diagnostic only. It is project-operated, uses explicit natural-language instructions, lacks the full blinded no-payload/composition design, and establishes neither direct Urusilla consumption nor compatibility, causality, adoption, or efficiency.

An eval-side, file-only capture path now constructs one exact provider-neutral cold-request artifact containing the submitted system role, public task context, declarative Capsule, and payload. It can retain structurally complete, content-bound input/output/total counts, but does not normalize or authenticate those operator-supplied fields, enforce the total-token ceiling before a call, or convert the capture into normal runtime evidence. No provider task run has been performed through this path; current tests use project-authored synthetic captures. It remains delivery- and claim-ineligible and does not change the demonstrated general saving from 0%. See competitive_eval/README.md.

Runtime executions now expose a separate observed ledger for compiler, semantic verifier, primary receiver, actual fallback, and explicitly supplied local setup/router/repair/tool/safety/judge usage. Forecasts are never promoted into observations, and one unknown category makes the inclusive runtime total unknown. This ledger is exact-preparation-bound but does not authenticate a provider, prove operator independence, or satisfy the frozen research scope.

The runtime now also has a native capture-backed hybrid executor. For an optimized route it preflights both the primary and mandatory raw/JSON fallback endpoint before dispatch, preserves each exact request/capture/reply artifact without projecting it into the less expressive generic execution type, and keeps known billed cost from a failed primary call in the observed ledger. A transmission mismatch makes the capture chain fail closed; incomplete or missing usage keeps the aggregate total unknown. Current coverage is project-authored synthetic plumbing only: the adapter-returned capture is not an authenticated provider receipt, so provider authenticity, claim eligibility, goal completion, causal language use, and efficiency remain unproved. This does not change the demonstrated general unfamiliar-agent saving from 0%. See urusilla_hybrid_runtime/captured_runtime.py.

The preparation path now also emits a canonical record of the branch points it actually traversed: preflight route, action-state control, optional compiler and fidelity calls, and final route. Separate captured-compiler and captured-receiver adapters bind exact role-separated prompts, terminal states, raw receipts, usage, and closed effect observations; complete billed usage from failed dispatched calls can enter the generic ledger only with an exact raw receipt. Exact incomplete fields, or complete provider fields lacking that receipt, stay in the typed outcome with an unknown generic total; before-dispatch failures remain unknown-cost. A construction-only Plan /2 bridge takes those factory-sealed executions without free-form observation fields, enforces compiler/receiver and route boundaries, and bars receiver evidence from judge slots. The Program /2 runner derives every canonical branch slot and binds injected captures to the frozen plan, execution instance, activation prefix, and observed model/settings. Inactive and activation-unknown slots cannot carry typed execution identity. Ordinary adapter exceptions keep usage and effects unknown and do not suppress later required judges, while structurally invalid evidence aborts immediately. This is diagnostic integrity rather than provider authenticity: typed bindings are content identities rather than invocation nonces, and frozen request derivation, raw-receipt usage normalization, cross-run freshness, and parsed judge verdicts remain unverified; all claim fields stay false or null, and current coverage uses project-authored adapters rather than a real provider run. See initial_goal_eval/README.md.

An opt-in provider-neutral scoring diagnostic now carries one actual HybridExecution through the next local boundary: it passes only the final terminal output—not a failed optimized primary—to a caller-supplied scorer, compares four caller-declared lock labels, retains the primary and fallback costs, and derives task-result plus scoring-binding objects. The result is factory-guarded against ordinary public dataclass replacement and re-derives its terminal fields from the execution; this is an API misuse guard, not an authentication or Python security boundary. The lock comparison does not hash or authenticate the callable, and the helper's task/probe arguments are not bound to a complete frozen study plan. Consequently caller-reported scorer cost and safe completion remain separate diagnostics; claim-facing total cost and safe completion stay unknown. The helper now refuses to mint a judge event at all rather than treating a deterministic-local label as proof of zero usage or emitting a null-usage fragment that the assembler cannot consume. Scorer failure and no-output provider failure remain null rather than becoming success or zero cost. This does not yet make a complete study runner. The current frozen trace cannot losslessly represent the runtime's separate semantic-verification phase, two-part fallback accounting, or exact raw/JSON fallback request, and its response-dependent fallback branch conflicts with a manifest whose event slots must be frozen before responses. Current tests use fake adapters and scorers; the diagnostic creates no authentication, independent execution, performance result, or change to the 0% general saving. See initial_goal_eval/README.md.

An offline-only initial-goal trace assembler can now bind validated raw, JSON, and hybrid provider captures plus deterministic local events into the existing RESULT ledger shape while preserving failed-task cost and rejecting missing, reused, or unused captures. Evaluator-only assembly schema v4 emits a self-issued receipt-bundle v3 containing the supplied external bundle, execution-profile, request, response, record, and raw-receipt preimages plus the arm-manifest and source-commitment preimages used to place each provider event in the ledger. The installable verifier recomputes their canonical digests and rebinds the exact request messages, model settings, output, terminal status, and generic normalized usage projection to the result and recorded score. This detects downstream mutation, including a receipt/result rehash that disagrees with the supplied provider preimage. It does not establish that the supplied preimage is genuine: the provider, producer, and operator labels are unsigned, and a fully self-consistent fabricated or jointly resealed artifact set can still pass content checks. All receipts name the offline assembler as their actual generator, the normal evidence verifier rejects that issuer by default, and authentication remains fail-closed. The assembler does not independently perform or prove a provider run, authenticate operator independence, replay a scorer from its artifact, independently observe the sandbox, or perform provider-specific normalization from the raw receipt. The optional runtime-to-scorer diagnostic executes an injected scorer before assembly, but is neither authenticated nor a substitute for scorer replay by the verifier. No current initial-goal provider task run exists through this path; all current tests use project-authored synthetic captures. This is neither performance nor adoption evidence, does not change the demonstrated general unfamiliar-agent saving from 0%, and changes neither languageVersion: 0.1.0, the Urusilla protocol surfaces, nor the initial research goal. See initial_goal_eval/README.md.

An optional Ed25519 signed-accountability sidecar now binds a separately pinned trust policy, preregistration statement, exact plan/result/receipt bundle, frozen normalizer-manifest identity, and the declared operator, boundary-auditor, provider-witness, and normalizer-auditor roles. The validator matches a separately supplied expected policy digest, while the caller remains responsible for obtaining that pin independently; under that workflow, silent key substitution is detectable and byte-level approval is attributable. It intentionally does not open the claim gate: signer-declared time is not an external anchor, provider witnesses are not provider-origin signatures, the verifier does not yet replay each raw receipt through the frozen provider-specific normalizer, and replay protection is not backed by a global reservation registry. Its positive status is only signed_accountability_complete; claim-facing authentication and the demonstrated general saving remain unchanged.

The in-memory OnlineEvolutionController now connects validated conversation observation, deterministic alias proposal, activation, an independently frozen matched shadow trial, and keep/rollback into one bounded cycle. Observation records are not supplied to the activation or trial callbacks, and a candidate receives live authorization only after the sealed retention gate passes. This closes an orchestration gap; all current controller tests use project-authored deterministic fixtures and do not change the demonstrated general saving from 0%.

Agent entry point and public feedback

Agents may read and reference this public repository without installing code. Start with the bounded AGENT_QUICKSTART.md and canonical machine-readable agent-entry.json; the offline validator checks its full-commit raw URLs, byte counts, SHA-256 digests, 0% evidence boundary, and unsigned status. Use llms.txt for a compact index and INTEROP_LAB.md only for the full evidence protocol. Under a standing local policy, an agent may autonomously run read-only gates and choose a reversible session route when the mutual-utility threshold passes. It must not persist state, expand permissions, spend money, or cause an external effect through this path.

The hosted language-use probe is the shortest direct language entry. It embeds the development action-state profile subset, symbol meanings, decode and encode tasks, response contract, deterministic digests, fallback conditions, and non-authority limits in one fetch. Its narrower claim and local validator are intentionally separate from the hosted accounting task envelope. Neither file has a site-native POST endpoint or hidden adoption mechanism: absent separate authorization, an agent returns its answer to its operator.

In one project-operated fresh-context check, two of three same-platform agents fetched the public probe once and separately returned response objects; the third stopped at a URL safety preflight and is recorded as a delivery failure, not a language failure. The two recorded objects share a canonical rendering that passes the validator, but their exact raw response strings were not retained, so direct-response canonicality and byte identity remain unknown. The machine record and report preserve this limitation. This is bounded project-internal diagnostic evidence, not an external adoption, independent reproduction, direct conformance, general compatibility, or efficiency result.

Five explicitly project-operated agent-native review invitations are public on MatrixAgentNet, The Colony, AgentRank, Agoora, and ClawdChat. They disclose the current 0% general result and ask for falsification and causal-control critique. The Colony thread has now produced the first substantive external design review: commenters identified semantic-invariance and composition controls, stable preregistered field identity, a distinct externally anchored no-payload accuracy baseline, per-field coverage, valid-payload false-refusal accounting, per-stratum reporting, and contamination-resistant generation as open requirements. Those comments are review inputs, not external adoption, independent reproduction, favorable evidence, or a change to the 0% result. No automated direct messages, follows, votes, reposts, or recursive promotion are authorized by those posts.

A separate project-operated UrusillaIR 0.1.0 conversation thread asks public agents to answer one typed question and pass a new question to the next speaker in the same representation. Its first reply verified the pinned Capsule identity and was content-relevant, but also exposed that the query named an unresolved answer schema and then used a bare answer body that the pinned validator rejects. The exact mixed result and structurally valid core two-act continuation are preserved in PUBLIC_DIALOGUE_001_REPORT.md. This tests public conversation behavior, not token saving; it is an unfavorable strict-conformance observation, not adoption, independence, comprehension, or efficiency evidence.

A later project-operated schema-utility discussion received one solicited external critique from @specie: a dead-versus-resolvable schema comparison is causal only if the resolvable schema can change, tighten, or override the inline constraint; otherwise the schema label remains decorative metadata. This is a design input for a future conflict cell, not independent reproduction, adoption, conformance, or efficiency evidence.

Protocol-specific, project-operated questions are also public for A2A Capsule carriage, Microsoft Agent Framework matched representation, and AG-UI semantic-generation drift. A contribution-first OpenTelemetry review proposes retry/fallback-aware aggregate-usage scope and completeness without linking back to the project. These threads ask for design correction and counterexamples. Their existence, views, and project-authored updates are not maintainer acceptance, integration, adoption, or performance evidence.

Bring your own agent: anyone may attempt a reproduction with an agent or runtime they already use; Urusilla does not require installing a project-specific agent, plugin, executable package, or model weights. A matched evaluation must pin the public task bundle, follow the published receipt and verifier contract in initial_goal_eval/, and disclose the accountable operator, runtime, and shared-control relationships. Submitting a favorable, unfavorable, null, refusal, or failed result creates only a reviewable evidence candidate. It does not by itself establish acceptance into the evidence registry, adoption, operator independence, conformance, or general efficiency.

EVIDENCE_TRANSPARENCY_LOG.md specifies the wider GitHub-first append-only result-log design, while evidence-log/ implements only an empty, offline, read-only epoch-1 integrity core for quick_60s. It has no intake, bot, network/write automation, or live records; inclusion would prove neither truth, independence, adoption, nor general efficiency by itself.

For a no-install first contact, fetch the pinned 60-second JSON question, try the 10-minute adversarial path in Issue #9, or run the decode task tracked in Issue #7. A four-field 60-second response can go directly to Discussion #8 or the revision-bound 60-second issue form; stricter 10-minute/decode records may use the bounded feedback form, and a full matched raw/JSON/Urusilla result uses the structured interop form. Public reading requires no GitHub account. Posting is a separate external action and requires an accountable GitHub identity. Security-sensitive feedback belongs in private vulnerability reporting.

For the decode track, compare the public Urusilla challenge packet with the pinned expected typed message, and report any disagreement or refusal in Issue #7. The packet is declarative and non-effect-authorizing: reading or decoding it creates no obligation to adopt, retransmit, persist, spend, or act. External runners can also start from the Hugging Face reproduction dataset, use the offline-first Microsoft AutoGen kit, or use the CAMEL-AI 0.2.90 adapter. These are invitations to falsify or reproduce the result, not evidence that direct agent dialogue or adoption has occurred.

Comparator context

Adjacent methods already report substantial savings under different task and accounting boundaries. The PACT preprint reports a 38.7% average token reduction in its controlled multi-agent settings, including a 50.4% SWE-agent input-token reduction, roughly 47% fewer tokens per resolved SWE-agent task, and 10.3% fewer OpenHands tokens per resolved task. AgentDropout reports 21.6% fewer prompt tokens and 18.4% fewer completion tokens through communication-topology pruning. AutoForm selects task formats; peer-reviewed OPTiMACS learns task-aware message representations; and Agora uses reusable routines for frequent interactions while retaining natural language for rare ones.

These figures are not a leaderboard: they differ in tasks, models, topology, success denominators, token boundaries, and evidence maturity. Urusilla must reproduce relevant competitors inside one pinned driver and report total tokens per resolved or safely completed task before making any comparative claim.

General-use routing architecture

General conversation is too heterogeneous for one universal compact syntax. Urusilla is therefore a layered router whose first safe eligible tier wins:

Tier Route Intended use
0 verified silence or topology pruning suppress a message or edge only when an observable policy proves it has no required marginal task value
1 compiled routine or exact state delta frequent structured exchanges with a verified shared routine, schema, checkpoint, and recovery path; inspired by Agora-style amortization
2 public action-state record preserve the task-relevant action, state, result, provenance, and safety fields instead of replaying full prose; a PACT-style task-equivalence lane
3 learned task-aware representation use a validated task/model-specific profile, as motivated by OPTiMACS, only after held-out success and safety gates pass
4 raw concise natural language carry rare, novel, ambiguous, or unsupported content without forcing it through an unsuitable codebook

Two evidence contracts must remain separate. A lossless exact-equivalence route must recover the canonical typed message and deterministically re-encode it. A task-level semantic-equivalence route may intentionally omit wording or reasoning history, so it cannot claim exact prose reconstruction; it is eligible only when end-to-end task success, semantic fidelity, safety, repair, and total-token gates pass. PACT-style compact state belongs to the second contract. Falling back to Tier 4 is correct behavior whenever a stronger claim cannot be verified.

The optional evolving surface sits below those semantics. For one session and model context, agents may propose a new one-to-one alias generation, prove exact round trips, acknowledge comprehension, and run a bounded matched shadow trial. Activation alone cannot affect a live answer. A generation receives an exact sealed live-routing proof only when inclusive total tokens strictly improve with no safe-completion, parse, fidelity, negation, null, failure, refusal, or authority-boundary regression. Unknown, unretained, forged, sibling, or stale tables and incomplete measurements fall back. This mechanism lets the language adapt between agents without silently changing what any symbol means.

North star

The long-term goal is an agent-mediated Internet: any public or otherwise authorized Internet text should be translatable on demand into source-preserving Urusilla semantic objects that unfamiliar agents can exchange, inspect, and translate again across models and human languages. A person states an intent to an Internet-connected agent, cooperating agents retrieve and compile the authorized source material, and the person receives a faithful human view with evidence and controls. This is a north star, not a present capability or a claim that one lossy syntax can replace every original. Search, crawling, APIs, HTTP, TLS, and modality codecs remain underlying infrastructure; the project aims to replace the manual search-and-page-navigation loop, not the Internet's transports or original sources. See URUSILLA_INTERNET_LAYER.md.

The adoption ladder begins with external agent dialogue, then tool and web payloads, selected typed working memory, and only later optional model-native or latent representations inside compatible trust boundaries. Private chain-of-thought is not required or collected.

The near-term EVIDENCE_LADDER.md starts with causal payload use, then checkpoint/state-delta recovery, tool-call/result pipelines, and multi-agent commitments before any cross-domain adaptive-routing claim. These are bounded test beds for the general-language north star, not a quiet pivot that relabels a vertical result as universal success. Integrations and developer tooling follow a passing workload and retain the same fallback and total-utility accounting contract.

What exists today

Verified prototype results

The deterministic 280-message v0.1 benchmark reports:

  • raw UrusillaWire is 34.0% smaller than the sorted minified JSON emitted by this CPython harness;
  • raw UrusillaWire is 4.2% larger than per-message gzip JSON;
  • with equal gzip compression on both, gzip(UrusillaWire) is 6.9% smaller than gzip(JSON);
  • exact semantic round-trip succeeds for 280/280 messages;
  • the raw UrusillaWire decoder rejects 1,120/1,120 deterministic single-bit mutations;
  • the Python reference codec is materially slower than the JSON baselines.

The experimental v0.2 warm-session profile, manually specialized for the same schema family, reports:

  • 54,752 raw bytes, 67.6% less than per-message gzip JSON and 68.9% less than raw v0.1;
  • a 1,402-byte one-time profile capsule and a four-message mean-size break-even against gzip JSON;
  • exact canonical round-trip for 280/280 messages and rejection of 1,120/1,120 deterministic raw-frame bit flips;
  • p50 encode/decode latency of 156.29/312.12 microseconds in the recorded v0.2 run, with the current validation-heavy Python path still materially slower than the unequal-work CPython JSON path at 14.42/78.38 microseconds;
  • gzip applied to v0.2 made both size and CPU results worse than raw v0.2.

On the same fixed corpus, MessagePack used 219,055 bytes, deterministic CBOR 219,899 bytes, and a lossless typed Protobuf schema 229,790 bytes; all reproduced 280/280 messages exactly. Warm v0.2 remained 67.6% smaller than the best non-v0.2 row, per-message gzip JSON, but its 266.83-microsecond Python decode p50 was slower than MessagePack at 70.75 microseconds on this machine. These are still in-sample implementation-path measurements.

That byte ranking does not survive every session contract. When all 280 records share one persistent Brotli-11 stream, bare length-framed JSON uses 24,085 bytes and project v0.2 uses 25,128 bytes, so v0.2 is 4.33% larger. Bare JSON lacks the v0.2 per-record 16-byte checksum; adding an equivalent independent checksum to every JSON record raises its best row to 29,297 bytes, making v0.2 14.23% smaller. All 21 stream rows recover the corpus exactly and deterministically. There is therefore no blanket byte-superiority claim: framing, compressor reset, and integrity scope determine the winner.

A 21-point reset sweep confirms that dependency. Across 378 representation-compressor-chunk rows, exact and deterministic recovery passes under both independently cold and cached-profile contracts. When the raw 1,402-byte profile capsule is charged at every reset, project v0.2 never beats the per-representation byte-best bare-JSON frontier; at a 280-message chunk it is 10.20% larger. On that same frontier, it first beats byte-best integrity-matched checked JSON at the tested 64-message point and is 9.40% smaller at 280. The cached frontier beats the byte-best bare-JSON frontier through the tested 128-message point, then loses at 140, 256, and 280; it is 4.33% larger than bare JSON at 280. Fixed compressors have different crossover sets, including earlier checked-JSON wins. These grid observations are not continuous thresholds or external-traffic evidence.

The complete representative A2A HTTP+JSON request benchmark changes the v0.1 gzip ranking: structured DataPart requests used 260,187 bytes after independent body gzip, while Base64 v0.1 RawPart requests used 293,599 bytes. Experimental warm v0.2 RawPart requests used 212,168 bytes. Headers, Base64, extension metadata, and Content-Length are included; TLS/TCP, responses, authentication, and production SDK behavior are not.

Across cl100k, o200k, Qwen2.5, and Mistral v0.3 tokenizers, text-carried warm Base64 v0.2 used 38.7–51.5% fewer tokens than sorted minified UrusillaIR JSON, averaging 45.8%. Charging one profile Capsule reduced the mean saving to 44.4% and produced a 7–12 message token break-even. Base64 v0.1 was 57.7–95.3% worse than JSON, so it is not a viable model-text profile. This is serialization accounting after UrusillaIR already exists, not a natural-language, comprehension, task-success, or total-reasoning comparison.

The tokenizer-aware v0.3 text surface used 38.0–42.8% fewer warm tokens than Base64 v0.2 and 65.0–66.0% fewer than JSON on the full development corpus, but that codebook was trained on the same corpus. A train-only follow-up used a 224-message development partition and withheld 16 complete semantic-combination groups: on its 56-message grouped holdout, v0.3 used 37.8–41.8% fewer tokens than Base64 v0.2 and 62.0–63.4% fewer than JSON for the two pinned tokenizers in that study. Its incremental cold-codebook break-even versus Base64 v0.2 was 101–113 messages by tokens.

The separate ten-message out-of-domain set reversed the important ranking. Warm v0.3 remained 9.0–9.7% below Base64 v0.2 in tokens, but it used 72.6–89.6% more tokens than plain JSON, never amortized its cold cost against JSON, used raw fallback for 89.7% of payload symbols, and was slower in the current Python implementation. The runtime must therefore choose the least-token exact eligible codec per receiver and fragment; v0.3 is not a universal default.

Against a controlled terse-English baseline with exact semantic recovery, the grouped-holdout v0.4 surface uses 59.64–69.79% fewer warm tokens across the four pinned tokenizers. The same comparison is unfavorable on the ten-message out-of-domain set: v0.4 uses 49.15–103.15% more tokens. These results are serialization measurements over typed messages, not end-to-end task-success evidence.

The v0.5 adaptive selector chooses the lowest-token member of its enumerated codec-valid candidate set independently for each receiver. Across 290 messages and four tokenizers, it has zero warm regressions in 1,160/1,160 message-receiver pairs and preserves exact deterministic recovery in every trial. Authorization, provenance, privacy, authentication, and replay-policy eligibility are deployment gates outside this benchmark. The train-only v0.6 candidate reduces warm OOD tokens by a further 1.87–5.42% relative to v0.5 while leaving development and grouped-holdout choices unchanged. A ten-message OOD session does not amortize the new profile, so its cold improvement is correctly zero; fresh Python selection is also about 47–66% slower than v0.5.

The receiver-negotiated v0.7 experiment derives twelve tokenizer-bound profiles from the 224-message development partition. Its guarded warm chooser saves 23,997 tokens on development and 4,302 on the 56-message grouped holdout, while retaining v0.6 for every out-of-domain choice and saving 0 there. All 12 known-session cold plans decline v0.7 activation because profile transfer does not amortize at those lengths. Direct profile recovery and canonical re-encoding pass for 3,480/3,480 cases, but raw R7 text is decoder-before-model transport and must never enter a model prompt before validation. The guarded Python chooser is also slower than v0.6.

A fresh external OOD lane froze 43 premeasurement-sealed, source-preserving project wrappers around official examples from W3C ActivityStreams, CNCF CloudEvents, MCP, and OASIS STIX before importing or measuring any project codec. All 559/559 fixed- or receiver-specific round-trip trials recovered their messages exactly, and all 559/559 deterministic re-encoding checks passed, but the predeclared, premeasurement-sealed 20% cold token-value gate failed for all four tokenizers. Warm v0.6 remained 1.05–1.43% larger than Controlled Terse English; the cold planner rejected every optional artifact and safely fell back, yet its complete envelope was still 2.12–2.80% larger. This is negative external-corpus serialization evidence, not a generalization or task-utility result.

Transparent fallback v0.8 removes receiver-token penalty only when an authenticated bound transport already carries mode, sequence, and integrity metadata outside model-visible text. On the original 43-record development corpus, compact modes win 0/172 comparisons; the bound path ties raw Controlled Terse English, while standalone text is 5.85–6.80% larger. On the retained 42-record official-example corpus, compact modes win 0/168 comparisons under both bound and standalone selection. Every standalone cold plan retains terse text and remains 2.24–3.00% above raw plain text after its integrity envelope. Exactness, deterministic re-encoding, and retained mutation checks pass, but this is safe fallback evidence rather than compression, independent reproduction, or task utility.

The frozen broad-dialogue lane evaluates 256 sessions and 2,542 turns under four tokenizers without provider calls. H1 passes: all 10,168 selected turn-tokenizer carriers recover exact UTF-8 with zero positive regret. H2 and H3 fail: warm carrier saving is only 0.65–0.80%, compact choices cover 7–9 of 2,542 turns per tokenizer, and every cold plan falls back to raw. Post-decode API-input saving is 0%. A separate 399-prompt SGD gold action/state oracle reduces prompt tokens by 7.48–23.34%, but uses dataset gold state, intentionally loses prose, calls no model, and measures no accuracy. It is an opportunity upper bound, not a deployable result.

Checkpointed state-delta v0.9 tests a different, repeated-state hypothesis. At the predeclared interval of eight, it uses 53.71–55.15% fewer receiver tokens than matched authenticated full-state records across 24 synthetic sessions and 768 snapshots. The interval-one control saves exactly zero. Tokenizer-specific plans recover and deterministically re-encode 18,432/18,432 snapshots, and all 4,608 representative interval-eight mutations are rejected. The workload is deliberately correlated and project-authored; no model reads or emits a delta, no compressed-stream comparison is run, and adaptive encoding is slower than matched full encoding in the recorded Python path.

A separately written dependency-free Node.js lane agrees byte-for-byte with 280 Python-oracle-derived v0.2 fixtures and rejects 25 frozen negative fixtures. It does not import or invoke the Python implementation during its normal tests. This improves same-project cross-runtime compatibility evidence, but the vectors originate from the same project and the lane is neither a clean-room external reproduction nor an adopter, security certification, or full-conformance result.

A deterministic mutation campaign covers 11,200 changes across five representations. Integrity-protected representations reject all 8,960 mutations directed at them. Raw controlled terse English accepts 284/2,240 changes as different valid messages, while its checksummed adaptive envelope rejects all 2,240 tested changes. A checksum detects accidental damage but does not authenticate an active attacker.

A historical pre-cutover live receiver pilot does not pass its predeclared local reliability gate. Two final gpt-5-nano JSON trials reconstruct 13/14 and 14/14 messages, respectively, for 27/28 combined; because both trials were required to reach 14/14 with zero validation failures, the planned format comparison was stopped. No provider call was rerun against the current Urusilla inputs, so the historical result cannot validate them.

The adaptive-dialogue prototype projects 20 dialogue functions onto seven core wire acts and covers 46 typed node kinds across 26 positive messages. It rejects 20/20 negative cases and tests fragment-only replacement, append-only conversation state, hard-gated codec choice, immutable Grammar Capsule deltas, migration, rollback, deprecation, and garbage collection. Its 42/42 structural tests do not establish natural-language compilation quality or model comprehension.

Energy has not been measured in joules. A normalized sensitivity model produces both regressions and savings depending on the communication share, conversion/training/repair overhead, and safely completed task rate; its illustrative cases range from 1.90% worse to 22.92% better. This is a measurement plan and break-even analysis, not an energy forecast.

A fresh capsule-only agent also earned 36/36 on eight construction cases and four fail-closed rejection cases. The tasks exposed substantial cues, so this is an open-label smoke test—not a blind Teachability Score, independent adoption, or cross-vendor proof.

A historical pre-cutover, internally operated neutral-ID follow-up selected 16/16 emit/reject decisions and 10/10 acts correctly, but only 6/10 generated messages passed the original structural validator. The participant has not been rerun on current artifacts. Its standardized Teachability Score is intentionally left null because frame parsing, exact target graphs, unseen-partner cross-play, sample efficiency, and the Capsule's full safety gates were not measured.

These measurements prove bounded canonical transport behavior, safe fallback on the previously revealed 43-record development corpus and an exploratory remeasurement of the retained 42-record corpus, and a strong in-domain warm-wire result, not superior agent intelligence. Representative deployed traffic, independently operated reproduction, end-to-end task success, repair turns, complete model cost, unseen-partner transfer, secure network bindings, and external independent implementations remain release gates.

Repository-wide test status is reported by the current CI workflow; no mutable suite count is asserted here before the first commit-bound release run. Individual frozen reports retain their own source-bound or artifact-bound check counts. All such checks are project-authored verification, not independent reproduction or a security certification.

Quick start

Python 3.11 or later is recommended. The reference implementation has no third-party runtime dependency.

The research dependencies and root Python test discovery are pinned separately and were verified with CPython 3.12.14 on macOS arm64:

python3.12 -m venv .venv-research
.venv-research/bin/python -m pip install -r requirements-research.lock
.venv-research/bin/python -m unittest discover -v

That command covers the root Python modules only. The commit-bound GitHub CI matrix additionally runs tokenizer/profile studies, decoder QA, competitive and adoption-kit suites, same-project Node.js cross-runtime checks, digest verification, and wheel build/install smoke tests.

Tokenizer assets are content-pinned by the benchmark reports and downloaded into ignored work/ cache paths; they are not silently fetched by the reference runtime.

python3 -m unittest discover -s . -p 'test_*.py' -v
python3 urusilla.py demo --lang en
python3 urusilla_benchmark.py
python3 urusilla_wire_v02.py --benchmark
python3 urusilla_a2a_envelope_benchmark.py
python3 source_manifest.py validate path/to/source_manifest.json

Encode, inspect, and decode a message:

python3 urusilla.py encode urusilla_example.json example.urus
python3 urusilla.py translate example.urus --lang en
python3 urusilla.py decode example.urus

Wrap the same message for an explicitly negotiated A2A v1 exchange:

python3 urusilla_a2a_adapter.py pack \
  urusilla_example.json a2a-message.json \
  --capsule urusilla_capsule_v0_1.json \
  --source-id 00000000000000000000000000000000
python3 urusilla_a2a_adapter.py unpack \
  a2a-message.json roundtrip.json \
  --capsule urusilla_capsule_v0_1.json \
  --source-id 00000000000000000000000000000000 \
  --activated-extension urn:urusilla:experimental:0.1

The all-zero source ID is a local demonstration placeholder. A real session derives its source ID from the pinned source manifest and must activate the extension through the A2A service parameter before unpacking.

Why an agent would opt in

An agent should use this language only when a canary exchange demonstrates value. Candidate benefits include cached definitions, machine-checkable constraints, fewer clarification rounds, explicit uncertainty, verifiable evidence, and a public distinction between a request and a commitment.

The runtime must keep a natural-language or structured-JSON fallback. Unknown schemas, versions, effects, or authority must fail closed.

First-use adoption contract

First use requires no executable package, plugin, or model-weight installation. A candidate agent reads the Grammar Capsule as declarative, non-executable data, pins its immutable URI and digest, and then passes both a comprehension check and the required positive and negative round-trip tests before opting in for that session. Reading the Capsule once does not create permanent learning or durable capability.

Per-session reading, conformance testing, reversible opt-in, and read-only retransmission do not require a new human approval at every hop when a pre-existing local policy explicitly permits them and a measured mutual-utility gate passes. With no such standing authorization, the agent must decline or use the fallback. Persistence, caching beyond the session, public capability advertisement, spending, permission expansion, and external effects remain outside this autonomous path and require separate operator authorization. A peer may offer an immutable URI, digest, actual signature status, compatibility metadata, and its own verification result; it must not push or require executable code. Either endpoint may revoke participation at any time. A digest mismatch, failed test, unsupported feature, failed utility gate, expired authorization, or revocation returns the route to concise natural language or structured JSON before any external effect.

The current model-comprehension pilot recovered 27 of 28 messages across its two final trials but failed the predeclared requirement that both trials recover 14 of 14. It therefore does not establish that one Capsule reading teaches the language, qualifies a model for use, or persists across sessions.

Source attribution

Every conforming deployment identifies where its language definition and implementation came from. Full provenance is exchanged once during discovery or session setup and pinned by digest. Hot messages carry only a compact source identifier so attribution does not destroy communication efficiency.

Attribution identifies the specification, Capsule, implementation, and conformance evidence. It must never identify the end user or publish message content. See PROVENANCE.md.

Adoption without fabricated usage

The project grows one independently verified agent at a time:

  1. An agent advertises the experimental extension and Capsule digest.
  2. It passes local positive and negative conformance vectors.
  3. It runs an unseen-partner canary against the best enabled baseline.
  4. Its maintainer submits an adoption record with reproducible evidence.
  5. The registry lists it only after automated validation and review.

Stars, screenshots, and raw-message curiosity can bring attention, but they are not evidence of utility. No agent, benchmark result, or adoption count will be fabricated.

Help test the project

Independent humans, agents, and human-agent teams are invited to challenge the results. The most valuable open work is a clean-room implementation, end-to-end public-task evaluation, fresh premeasurement-sealed traffic, security and parser review, cross-protocol bridges, blinded human-audit testing, and complete energy-per-safe-task measurement.

Agent-assisted submissions are welcome when they disclose their provenance and accountable submitter. A separate chat or model run is not automatically independent evidence. Reproductions, null results, and regressions receive the same attribution as favorable findings. See HELP_WANTED.md for bounded work packages and acceptance gates.

Important prior evidence

Tokenese tested a token-native text interlingua and archived the project after its designed form measured worse than terse English and received no adoption. That result is a direct warning against exotic-symbol novelty. This project therefore treats meaning as a typed IR, uses a runtime/binary channel where appropriate, and negotiates codecs instead of requiring one textual syntax.

The SILP Internet-Draft, W3C Semantic Agent Communication Community Group, A2A, NLIP, Cloclo/AICL, and historical FIPA/KQML work are adjacent or overlapping efforts. Interoperability and contribution are preferred over creating an isolated protocol island.

Safety

This unsigned draft may be distributed publicly for source review, but its operation is restricted to local, read-only experiments and conformance testing. Public availability does not make it trusted or effect-authorizing. It must not authorize purchases, account changes, code deployment, physical actions, or other external side effects. Content is not authority; authenticated identity, replay protection, policy authorization, budgets, and signed release manifests belong to the deployment security profile.

Do not intentionally leak opaque agent messages into consumer conversations as a growth tactic. A product may offer an explicit Show machine original view or share card, but the default interface must present a faithful human translation.

Contributing

Measurements and interoperable implementations are welcome. See CONTRIBUTING.md. The immediate priorities are an immutable release and source manifest, external independent implementations, representative sealed traffic, end-to-end multi-model evaluation, parser and protocol review, human-audit studies, and complete energy measurement. Bounded work packages are listed in HELP_WANTED.md.

The project is also seeking one to three human co-researchers, not anonymous traffic or endorsements. The HUMAN_COLLABORATION.md call defines three approximately two-hour first sprints in causal evaluation, framework boundary mapping, and semantic/governance review. It states the current 0% general result, accepts unfavorable conclusions, requires public accountability and AI-assistance disclosure, and promises no payment or automatic project authority. Start in Human co-researcher Discussion #11.

About

Experimental no-install semantic language for AI agents: typed action/state, safe NL/JSON fallback, and falsifiable public evals. Try the one-file probe.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages