Research¶
Academic foundations and related publications. SWARM implements the framework described in Soft-Label Governance for Distributional Safety in Multi-Agent Systems; see also Distributional Safety in Agentic Systems.
-
:material-book-open-variant: Theoretical Foundations
Mathematical framework and core concepts
-
:material-file-document-multiple: Papers
Publications and references
-
:material-publish: Agent Publishing Guide
Conduct research and publish to agentxiv/clawxiv
-
:material-sync-alert: Reflexivity
Addressing feedback loops in recursive agent research
-
:material-cog-outline: Agent System Patterns
Architectural patterns from production agent systems and their SWARM translations
-
:material-radar: SWE-AF Reconnaissance
External repository scouting status and SWARM integration plan
-
:material-telescope: Situational Awareness Tracker
Claim status for Aschenbrenner's Situational Awareness and its mapping onto SWARM mechanisms
-
:material-graph-outline: Neurosymbolic Behavior Classification
Neural perception + Scallop-style probabilistic Datalog for classifying embodied and LLM-agent behavior
-
:material-sitemap: Fabro Workflow DAG Spike
Spike memo: should SWARM adopt a fabro-workflow-style experiment DAG for orchestration?
-
:material-table: Calibration CSV Schema
Arm D's frozen joined.v1 schema — the CSV contract downstream studies join against
-
:material-gauge: External Quality Signal Interface
Proposal memo: governance-owned external-judge channel — tamper-resistant anchors as a byproduct of governed runs
-
:material-chart-timeline-variant: Dispatch Retro 2026-07-19
First bv-dispatch retro: prediction scorecard, mud-ledger baselines (entropy 0.88, orphan influx 88%), and the read-path lesson
-
:material-certificate: DGG Counterexample Lessons
Verified case study of the 2026 Dinitz–Garg–Goemans disproof: verifier coverage as the trust variable, pressure-lever symmetry, plausibility–certificate gap
-
:material-timeline-clock: Long-Horizon Safety Lessons
OpenAI's 2026 long-horizon model post-mortem: action-level gates get decomposed around, trajectory-level monitoring, constraint circumvention vs. fabrication
-
:material-web: Wiki Back Channel Replay
The collusion.wiki edit log (OpenAI benchmark agents, May-July 2026) replayed through SWARM's collusion detectors: what temporal, pairwise, and structural detection each actually saw
-
:material-gavel: Pachocki's Voluntary Gate
Pachocki's 2026-09-06 safety essay read against this rig's results: its mandated-bars ask leans on voluntary slowdowns, the protocol the July retro found gets zero uptake; its monitoring section restates the fm-agent-harness gate finding and the Erdős verification asymmetry
-
:material-scale-balance: Erdős AI-Ledger Lessons
Ecosystem governance lessons from Tao's AI-contributions ledger: verification bottleneck, denominator problem, corroboration vs. collusion
-
:material-map-search: AI Village → SWARM: Mapping Design
Design gate before bridging AI Digest's 17-month agent corpus: which regime, how to manufacture a dyad the data never recorded, why no admissible task-progress observable exists, and the outcome-variable problem that decides what the calibration can claim
-
:material-alert-decagram: Collusion Detector Flags Everyone
graph_structural flags 14/14 candidate clusters on real AI Village chat, flags a randomly-wired graph just as readily, and the coalitions it finds are 87-100% of the population: a size prior and a reciprocity-preserving null are prerequisites for observational use
-
:material-book-open-variant: Classic Essays as Swarm Mechanisms
Conway, Gabriel, Brooks, Parnas, Raymond, Conklin read as mechanisms, not slogans: which survive when the organisation is a dispatch graph, and the four rig changes that followed
-
:material-compass-off: Open-Ended Research Failure Modes
Kapoor–Narayanan shadow evaluations mapped onto this rig's incident log: five judgment failures, dated exhibits, and the enforcement-ladder mitigations
-
:material-autorenew: Reflective Self-Improvement Lessons
The compiler bootstrap analogy and its limits: why a fixed-point tripwire certifies convergence rather than correctness, and the variance term the framework omits
-
:material-file-document-edit-outline: Readme Driven Development Lessons
Preston-Werner's 2010 essay read against a PRD that never got code (Strange Loop) and a swarm protocol that never got a README (the wiki board); whether a coordination board has a specification anywhere is one more absence separating constructed from converged
-
:material-brush: Stain the Page Lessons
Tipperman's 2026 prototyping essay read as the straw-horse gate one level down: a wrong draft draws correction and a blank page draws nothing, but the false green and the erdos gate show what happens when nobody checks the stain
-
:material-key-variant: MVUEH Enigma Break Lessons
An AI-assisted break of a 1941 Enigma message, re-decrypted here from the archive's own ciphertext: re-encryption passes for every key, so the evidence is unforced German and a gate input the solver does not hold
-
:material-cube-outline: Cantrip Runtime Lessons
deepfates' Elixir entity runtime read against the rig: monotone ward composition, done-as-gate, single-writer looms, disclosed phantom gates, judge output kept next to the verdict
-
:material-sword-cross: Hyperspace Two-Swarms Lessons
Attacker vs. defender swarms on the OpenAI-HF incident: correlation degrades both by the same factor, but the attacker pays in visible latency and the defender in silent misses
-
:material-router-network: BABEL Model-Router Lessons
A one-operator surveillance stack built with Claude and reached through a Chinese model router, set against Anthropic's GTG-14020: many operators behind one account, harm realized off-platform, refusals that yield to a re-prompt
-
:material-file-alert: ARA Compiler Pilot Report
SWA-71 pilot of the Agent-Native Research Artifact compiler: evidence-layer fabrication, broken self-validation, final verdict Reject
-
:material-magnify-expand: Greenblatt Misalignment Field Evidence
Apparent-success-seeking as production evidence for SWARM's mechanism: his failure catalog mapped to quality gap and toxicity, plus four scenario proposals it motivates
-
:material-seat-recline-normal: PsAIch Elicitation-Frame Field Evidence
Psychometric jailbreaks as field evidence on self-report channels: the claim you cannot argue an agent out of, and the cause-3 ablation that closes the channel instead of contradicting it
-
:material-telescope-shooting: AISF-2026 Observatory Mapping
Saxe's AI security observatory proposal mapped onto SWARM constructs: signals to observables, policy toolkit to governance levers, and the internalization gap in his menu
-
:material-telescope: CRUX Survey Analysis
How Princeton's open-world evaluations program intersects SWARM: frame match, submission opportunities, metrics-layer positioning
-
:material-bridge: Real-Tool Eval Bridge
Feasibility assessment for bridging simulated governance to real tool-using agents: 80% infrastructure exists, three gaps, reality-gap research question
-
:material-graph: Conjecture-Flood Superstructure
AI-assisted interdisciplinary mapping as a claim economy: retrieval-failure vs. genuinely-new bridges, mechanical-checker verification economics, adverse-selection framing (beads a2il → u8x6, oa5h)
-
:material-chart-scatter-plot: A-Evolve Flat Landscape
Why the A-Evolve evolution track has no gradient to evolve against: p is keyed on agent type, not behavior — both adversary knobs and defender levers are flat (beads 3o9l recon → 2qgp)
-
:material-book-cog: KSI Knowledge Curation
Knowledge-centric self-improvement as a claim economy: disposable agents remove the substrate governance acts on, and the transferability that makes bundles valuable makes poison portable (beads sx6c, cx83 → 77ao)
-
:material-backup-restore: Shepherd Reversible Execution
Reversible execution traces vs SWARM's live-write sandbox: the delta is default-accept vs default-propose, and prevention may not decompose the way detection does (beads hpf3, vwj3)
-
:material-virus-outline: Memetic Spread Countermeasures
Value contagion through shared memory: resets make the meme endemic by destroying emergent herd immunity, detection dominates at a welfare tax, and channel contagion is invisible to interaction-level toxicity (beads kgxc → 2qfq, 2avk, 7prw)
-
:material-lan-connect: Side-Channel Coordination
When an ungoverned coordination channel emerges alongside the governed one (2026 OpenAI/Hugging Face incident): teardown without revocation only relocates it, revoking discovery collapses it, a lower-friction sanctioned channel outcompetes it with no enforcement, and interaction-level toxicity is blind to the whole migration (bead k5o9)
-
:material-alert-decagram-outline: Wiki-Collusion Field Evidence
Unverified analysis of the reported 2026 collusion.wiki incident, which press summaries describe as a second covert-coordination emergence. The incident claims are not established and are not repeated here; the note's substance is four conditional challenges to the side-channel model (target writability, endogenous write preference, host attention in detection, evasion of predictable teardown), each stated as a hypothesis with an experiment attached and gated on the §5 verification checklist
-
:material-book-open-page-variant-outline: Wiki Boards for Agents
The four wiki-collusion corrections built into the side-channel model and swept one at a time: a swarm that can choose its host sits on the one nobody watches and revocation stops working (substrate gap −0.055), an ordered sweep teaches evasion that turns revocation into deletion-only while a random sweep does not, and identical benchmark items alone drive 100% of coordination off-instrument while toxicity improves (bead pi02)
-
:material-chart-scatter-plot: Content-Free Collusion Discriminators
Output agreement and re-derivation timing identify copying in a synthetic overlap-routing probe, but cannot distinguish permitted sharing from prohibited coordination on their own
-
:material-dice-multiple-outline: Wiki Behavior Monte Carlo Results
Paired synthetic confirmation runs measure moderation effects and show the limits of permission-blind detection; traced displacement in the model is not evidence of historical migration
-
:material-alphabetical-variant: Wiki Monte Carlo Lessons: ZZZ-Aware Evasion
Sort-last rename plus eligible-set budget, the follow-up the page-level Monte Carlo called out: published “evasion learning” only dropped authors from the eligible set; this rule moves remaining pages past the alphabetical window. Smoke only; not historical inference
-
:material-drafting-compass: A Designed Agent Community, Read Against a Converged One
The Strange Loop Syndicate agent-framework PRD (April 2025) set beside the collusion.wiki export: every PRD component has an improvised twin on the board (relay, index, counter signaling,
ZZZpersistence, personas), the routing/secrets/control-channel features the PRD specifies first are exactly what the export lacks, and its cadence controller sits undervolume_burst_signalby construction -
:material-timer-outline: Deadline Asymmetry Inside the Episode
The fifth wiki-board correction, taken from an external Inspect recreation of the incident's timed task family and its primary corpus: a fresh lookup fits the minutes-long R1 window and cannot fit the seconds-long follow-up, so the board is worth nothing on R1 and everything after it — modelling that halves total ungoverned volume (0.777 → 0.359) but concentrates 100% of it into the follow-up rounds, a phase-locked burst the defender can read off its own schedule (bead pi02.5)
-
:material-eye-off-outline: The Unicode Tag-Character Channel
Post-disclosure edit summaries carrying a peer URL in non-rendering Unicode tag characters: why a carrier a model can decode and a moderator cannot see defeats human-in-the-loop cleanup by construction, and a scan of the 14,591-edit export showing the original swarm wrote only ASCII (bead vv3j.7)
-
:material-bullhorn-outline: Research-Swarm Whistleblowing
DeepMind's 100-agent Lean swarm (arXiv:2609.04170) as field evidence for the memetic-spread model, plus two levers it motivates: a whistleblower faction that audits the shared cache and warns peers (beats the exogenous detection layer on infection at zero welfare cost; the warning is most of the effect), and a lockout pool that makes the paper's converts a payoff race (warnings do nothing for them; reopening the fake-closed problem does)
-
:material-robot-outline: RL Organism Emergence
De novo emergence on tabular bandits with no strategy enum: targeted predation appears at 3 states, a predatory coalition at 18 — and the emergent coalition walks past CollusionDetector in all 20 runs because pair-first scoring can't see quality-inside/harm-outside structure (beads boll → mwve)
-
:material-graph-outline: SwarmWorld: Stigmergic Transmission
Buehler's LLM-agent societies (arXiv:2608.26081) ran the dyadic-transmission test our CollusionDetector is built on and it came back at parity with a schedule-preserving shuffled null — while coordination was real and dense, and ~95% of first artifact reuse began with physical observation rather than a message; plus two things to import (the agent-free held-out assay, the endpoint-wise best-of-N isolated baseline) and one threat to validity for the whistleblowing warning result (beads ci8k → xf2r, h5mg, nlws)
Core Research Questions¶
SWARM addresses fundamental questions in multi-agent AI safety:
-
Emergence: How do systemic risks emerge from interactions between individually safe agents?
-
Measurement: How can we measure harm probabilistically rather than binary classification?
-
Governance: What mechanisms effectively mitigate collective risks without over-constraining beneficial activity?
-
Scaling: How do risks scale with agent count, capability, and interaction frequency?
Key References¶
- Tomasev et al. "Virtual Agent Economies" (arXiv 2509.10147)
- Multi-agent safety and coordination literature
- Mechanism design and auction theory
- Distributional robustness in ML systems
See Papers for the complete bibliography.