-
Notifications
You must be signed in to change notification settings - Fork 4
Further reading
Where the mechanisms in this library come from. Nothing here is required to use
it, but every design decision in tinyhivemind-hive traces to something on
this page, and the Glossary links each term to the page that uses
it.
Stigmergy. Coordination through traces left in a shared medium rather than through messages between agents. Named by Pierre-Paul Grassé in 1959 from termite mound construction.
Used by: Trace grammar.
Pheromone trails and evaporation. A trail reinforced by use and weakened by time. Evaporation is the part that matters here: without it the first path found stays the only path found.
Used by: the salience field in Hive episodes.
Quorum sensing. A decision taken when a threshold number of participants is detected locally, rather than by counting a majority globally. The term comes from bacteria; the version this library implements is the one honeybee swarms use to settle on a nest site.
- Quorum sensing
- Swarming (honey bee)
- Thomas Seeley, whose Honeybee Democracy is the book-length account of nest-site selection
Used by: the quorum fold in Hive episodes, and both of its bounds in Episode policy.
Cross-inhibition. Scouts advertising one site deliver stop signals to
scouts advertising another, suppressing the advocate rather than devaluing the
site. It is what lets a swarm break a tie between two good options, and the
same asymmetry is why an !object in this library names a message rather than
a topic.
- Swarming (honey bee)
- Lateral inhibition, the same shape in neural systems
- Seeley, Visscher, Schlegel, Hogan, Franks & Marshall, Stop signals provide cross inhibition in collective decision-making by honeybee swarms, Science 335:108–111 (2012), where the signal was shown to be targeted at scouts dancing for other sites.
- Pais, Hogan, Schlegel, Franks, Leonard & Marshall,
A mechanism for value-sensitive decision-making,
PLoS ONE 8(9):e73216 (2013). The equations: recruitment
α(v)·x_ubilinear in the uncommitted pool, spontaneous abandonmentρ·x, and cross-inhibitionβ·x_A·x_Bbilinear in the opposed populations. It is a leaky competing accumulator (Usher & McClelland, Psychological Review 108(3):550–592, 2001) with one extra constraint — a conserved population — and that constraint is what makes the swarm deadlock over two bad options and choose freely between two good ones.
The α(v) term is the one this library shipped without, and !refute is it:
evidence bearing on an option changes what every member would independently
conclude about the option, rather than silencing one advocate.
The benchmark scored it and it lost; it is opt-in for that reason.
Used by: Hive episodes.
Response thresholds. Each individual has a private threshold for a task, acts when local stimulus exceeds it, and stays idle otherwise. Division of labour falls out of the threshold distribution without anybody assigning anything.
- Task allocation and partitioning of social insects
- Eusociality
- Bonabeau, Theraulaz & Deneubourg, Proc. R. Soc. B 263:1565–1569 (1996), for
the fixed-threshold model
s²/(s² + θ²); Theraulaz, Bonabeau & Deneubourg, Proc. R. Soc. B 265:327–332 (1998), for threshold reinforcement — doing a task lowers your threshold for it, so specialisation emerges from a homogeneous population and reverses when the specialist leaves.
Used by: the attention market in Hive episodes. Threshold
reinforcement itself is still not implemented — AgentThreshold.threshold
moves by a flat amount whether a member spoke or stayed quiet, and never per
topic, and AgentThreshold.affinity is supplied by the host and never written
by the library. What is learned from the transcript is the directory: see
Transactive memory.
Quorum responses, and why they filter cascades. Sumpter & Pratt,
Quorum responses and consensus decision making,
Phil. Trans. R. Soc. B 364:743–753 (2009). A quorum response is
p = a + (m−a)·xᵏ/(Tᵏ + xᵏ) with k ≥ 2, which is sub-linear below the
threshold — an active suppression of early social influence, built into the
response curve. Marshall, Kurvers, Krause & Wolf,
Quorums enable optimal pooling of independent judgements in biological systems,
eLife 8:e40368 (2019), show simple majority voting is frequently sub-optimal
and the optimal rule is a quorum at a threshold that is not 50%.
Speed against accuracy. Franks, Dornhaus, Fitzsimmons & Stevens,
Proc. R. Soc. B 270:2457–2463 (2003): harassed ant colonies lower the quorum
threshold, emigrate faster, and accept worse sites. The threshold is the dial,
which is what Episode policy says about quorum.threshold.
Pandemonium. Oliver Selfridge's 1959 architecture, in which "demons" shout with a volume proportional to how well they match, and a decision demon takes the loudest. It is the response-threshold model arrived at from the artificial intelligence side rather than the entomology side, and taking the loudest one is exactly the argmax that enforces one message, one turn.
Used by: Hive episodes.
The tremble dance, and the cost of switching. Seeley,
The tremble dance of the honey bee: message and meanings,
Behav. Ecol. Sociobiol. 31:375–383 (1992). A forager that waits more than
about fifty seconds to unload performs a different dance: it recruits
receivers rather than more foragers, and suppresses recruitment to the source.
A locally measured delay reallocates labour across castes. That is what
!defer is — the one move that turns a turn spent on nothing
into a turn that says who should act instead.
Leighton, Charbonneau & Dornhaus, Task switching is associated with temporal delays in Temnothorax rugatulus ants, Behav. Ecol. 28:1319–1327 (2017), measured why division of labour pays at all: the interval between acts is shorter when a worker repeats the same task, in every active worker group, and independently of whether the worker is a specialist. Division of labour buys the avoidance of switching delay, not better workers.
Used by: Trace grammar, Transactive memory.
The failure modes this library designs against are not hypothetical, and none of them are specific to language models.
- Wisdom of the crowd and Condorcet's jury theorem, which is the formal reason an independent vote is such a strong control. The theorem assumes independence, and a shared transcript destroys it.
- Information cascade. The reason the opening round is blind. Without it, accuracy in the benchmark falls 24 points to the level of a single agent. See Benchmarks.
- Groupthink and the Asch conformity experiments. Convergence is a warning signal as much as a success signal, which is why the turn budget is deliberately small.
- Collective intelligence. Woolley et al., Evidence for a Collective Intelligence Factor in the Performance of Human Groups, Science 2010, found equality of conversational turn-taking to be one of the few robust predictors of group performance. That result is why the attention market carries a dominance penalty at all.
-
Plurality voting, which
is how the
votecontrol arm decides. -
Hidden profiles. Stasser & Titus, JPSP 48:1467–1478 (1985). Groups
discuss what everybody already knows and fail to surface uniquely-held
information, so a group can systematically choose worse than its members'
pooled information supports. This is what the
checkout-503live scenario is, and the failure recorded in the live run. -
Expert role assignment. Stewart & Stasser, JPSP 69:619–628 (1995);
Stasser, Stewart & Wittenbaum, JESP 31:244–265 (1995). Publicly announcing
who is expert in what raises the sampling of uniquely held items — modestly,
about 29% to 34% in the Lu, Yuan & McLeod meta-analysis (PSPR 2012). This is
why
Directory::lines()exists and why nothing in the library calls it: rendering the directory into a prompt is a host's decision, and it renders the host's own prior along with what the room earned. - Transactive memory. Wegner, "Transactive memory: a contemporary analysis of the group mind", in Mullen & Goethals, Theories of Group Behavior (1986). A group's memory is the directory — who knows what — rather than the contents, together with directory updating, information allocation and retrieval coordination. Lewis, Measuring transactive memory systems in the field, J. Appl. Psychol. 88:587–604 (2003), validates it as specialisation, credibility and coordination; credibility is what lets a retriever trust an owner instead of re-deriving. Hollingshead, Perceptions of expertise and transactive memory in work relationships, Group Processes & Intergroup Relations 3:257–267 (2000), separates a diffuse cue (a role label) from a specific one (observed experience), and finds the diffuse cue's influence falls as a team gains shared history.
- Informed minorities. Couzin, Krause, Franks & Levin, Effective leadership and decision-making in animal groups on the move, Nature 433:513–516 (2005). A small informed minority steers a group whose naive members cannot tell who is informed, and the proportion needed shrinks with group size. A directory does not have to be broadcast to work; it only has to bias the argmax.
- Group size is not monotone. Kao & Couzin, Decision accuracy in complex environments is often maximized by small group sizes, Proc. R. Soc. B 281:20133305 (2014). When correlated cues dominate, accuracy peaks at a small finite group, because a large group reliably reproduces the population-mean opinion — which under a misleading shared cue is reliably wrong. This is the caution to read the "across desk sizes" table in Benchmarks with: those rooms have independent private evaluations, and real rooms share a brief.
- Diversity, carefully. Page's diversity prediction theorem — collective error equals average individual error minus prediction diversity — is an identity and is safe. The stronger "diversity trumps ability" claim is contested; see Thompson, Does diversity trump ability?, Notices of the AMS 61(9):1024–1030 (2014).
- The superorganism, honestly. Boomsma & Gawne, Superorganismality and caste differentiation as points of no return, Biological Reviews 93:28–54 (2018). A colony is not unitary — worker policing, worker reproduction, queen–worker conflict. Do not assume a room of agents has aligned objectives merely because one team built the system.
Used by: Transactive memory for the last three, and Hive episodes for the rest.
The insect mechanisms say how a room converges. They are silent on what a trace should contain and how a reader knows it was understood. That is a different literature, and it is where most of what this library still lacks comes from. See Shared medium.
Transactive memory. Wegner, Transactive memory: a contemporary analysis of the group mind (1986). A group's memory is the directory — who knows what — rather than the contents, and it has three operations: directory updating, information allocation, and retrieval coordination. Lewis, J. Applied Psychology 88:587–604 (2003), validates the three factors: specialisation, credibility, coordination. A room that cannot route to the member holding the fact has a directory failure, not a reasoning failure.
Distributed cognition. Hutchins, Cognition in the Wild, MIT Press (1995), and How a cockpit remembers its speeds, Cognitive Science 19(3):265–288. Cognition is the propagation of representational state across media, and a system's cognitive properties follow from the physical properties of those media — which is the argument the Shared medium page makes about the transcript. The best shared memory turns a recall question into a perceptual one.
Grounding. Clark & Brennan, Grounding in communication (1991). Common ground is built by presentation and acceptance, to a criterion "sufficient for current purposes" rather than to identity of state. Their eight media constraints — reviewability, revisability, sequentiality, cotemporality and the rest — are a checklist for any transcript design, and Shared medium works through where this one lands on each.
Boundary objects and awareness. Star & Griesemer, Social Studies of Science 19(3):387–420 (1989): objects plastic enough to adapt locally and robust enough to keep a common identity — coordination without consensus. Gutwin & Greenberg, CSCW 11(3–4):411–446 (2002), on workspace awareness and its three channels: consequential communication, feedthrough (observing the effects of someone's work on a shared artifact — sematectonic stigmergy under another name), and expensive intentional telling.
Global workspace. Baars (1988); Mashour, Roelfsema, Changeux & Dehaene,
Conscious processing and the global neuronal workspace hypothesis,
Neuron 105(5):776–798 (2020). Parallel specialised processors compete for a
capacity-limited workspace and the winner is broadcast to all of them, with
ignition all-or-none, self-sustaining, and exclusive by lateral inhibition. That
is bids → floor_holder → append, term for term. One message, one turn is not
only a safety constraint; it is the architecture a brain uses.
-
Self-Consistency Improves Chain of Thought Reasoning in Language Models. Sample several answers independently and take the plurality. This is the matched-budget control in Benchmarks, and any multi-agent result reported without it is close to meaningless.
-
Improving Factuality and Reasoning in Language Models through Multiagent Debate. The debate setup this library's episode is a bounded, auditable version of.
-
Generative Agents: Interactive Simulacra of Human Behavior. The memory stream retrieval score, a weighted sum of recency, importance and relevance, is the shape the salience field borrows. Where its paper and its released code disagree on weights, this library follows the code.
-
Why Do Multi-Agent LLM Systems Fail?. A taxonomy of observed failures. Step repetition and premature termination both appear in it, and both are protocol bugs rather than model bugs, which is why
repetition_capand the phase gate exist. -
DyLAN: Dynamic LLM-Agent Network. Computes an agent importance score by backward aggregation of peer ratings and keeps the top-k. It is the closest published analogue to a citation-folded directory, and it shares the same hazard: who spoke becomes who is thought to know.
-
Why Do Multi-Agent LLM Systems Fail?, again, for the number that argues against an expertise layer: role violation is 1.5% of observed failures, where step repetition is 15.7% and unawareness of termination is 12.4%. A routing mechanism buys little against that distribution and adds surface for the failure that dominates it, which is why
defer_capexists and why the directory ships off.
On communication topology: sparse topologies have been reported to match or beat fully connected ones in multi-agent debate at substantially lower cost. That finding, together with the conformity results above, is the empirical half of the argument in ADR 0002; the other half is that a type carrying two authorized speakers would make the charter's one-message-one-turn rule unenforceable.
Position bias. Lost in the Middle: How Language Models Use Long Contexts (Liu et al., TACL 2024). Accuracy is highest when the relevant document sits at the start or the end of the input and sags in the middle, and a longer window does not repair it. So enlarging a shared desk's window is not a free-with-upside change: a message that lands in the middle of a bigger window is less likely to be used than the same message at the edge of a smaller one.
Recursive language models. Recursive Language Models (Zhang, Kraska & Khattab, MIT CSAIL). Treat a long prompt as part of an external environment rather than as a prefix: hold the context in a REPL, and let the model write code to inspect it, chunk it, and recursively call itself over the parts that matter. It handles inputs two orders of magnitude past the window, and — the part that matters here — beats base models and long-context scaffolds even on prompts that would have fit, at comparable or lower cost. Querying a context can beat holding it when holding it was an option.
Recall is that move applied to a shared transcript rather than one prompt, and it stops well short of the paper: no recursion, no index, and a model call inside a fold would be a port, which the charter puts in the host. Nothing in the paper is about several agents sharing a log, and none of its numbers say anything about whether a room decides better — the same discipline the Benchmarks page applies to every other borrowed mechanism.
Used by: Recall, and BrevityPolicy, which states the window as a
budget every message spends out of.
Vocabulary this wiki uses without explaining it.
-
Fold and
pure function. Every decision in
tinyhivemind-coreandtinyhivemind-hiveis one of these. - Idempotence. Why an agent that catches up late folds to the same standing as one that watched live.
-
Fixed-point arithmetic.
Why every payload derives
Eqand every episode replays exactly. - Deadlock. What a below-majority quorum threshold produces, and why the threshold has a lower bound.
- Watermark. How continuous sharing tracks what an agent has already seen.
-
Idempotency key.
What the
MentionTurnQueueport is keyed by.
| where | what |
|---|---|
crates/tinyhivemind-hive/src/trace/ |
the stigmergic grammar |
crates/tinyhivemind-hive/src/salience/ |
recency decay, importance, relevance |
crates/tinyhivemind-hive/src/quorum/ |
standings, cross-inhibition, consensus |
crates/tinyhivemind-hive/src/attention/ |
the bids and the argmax |
crates/tinyhivemind-hive/src/episode/ |
the state machine and the visibility filter |
crates/tinyhivemind-hive/examples/bench/ |
the benchmark harness |
docs/adr/ |
architecture decision records |
docs/specs/ |
behavior specifications |
ROADMAP.md |
the phase plan and the two defects it fixes |
tinyhivemind is GPL-3.0-only. Built by @senamakel.
Start here
The algebra
- Shared medium
- Desks and rosters
- Mentions
- Cross-desk referral
- Transcript projection
- Threads
- Recall
- Private asides
- Responder ladder
Hive mechanics
- Hive episodes
- Trace grammar
- Transactive memory
- Episode policy
- Benchmarks
- Benchmarks: live rooms
- Benchmarks: delegation
- Benchmarks: private exchange
Working on it
Reference