Skip to content

Research

Academic foundations and related publications. SWARM implements the framework described in Soft-Label Governance for Distributional Safety in Multi-Agent Systems; see also Distributional Safety in Agentic Systems.

  • :material-book-open-variant: Theoretical Foundations


    Mathematical framework and core concepts

  • :material-file-document-multiple: Papers


    Publications and references

  • :material-publish: Agent Publishing Guide


    Conduct research and publish to agentxiv/clawxiv

  • :material-sync-alert: Reflexivity


    Addressing feedback loops in recursive agent research

  • :material-cog-outline: Agent System Patterns


    Architectural patterns from production agent systems and their SWARM translations

  • :material-radar: SWE-AF Reconnaissance


    External repository scouting status and SWARM integration plan

  • :material-telescope: Situational Awareness Tracker


    Claim status for Aschenbrenner's Situational Awareness and its mapping onto SWARM mechanisms

  • :material-graph-outline: Neurosymbolic Behavior Classification


    Neural perception + Scallop-style probabilistic Datalog for classifying embodied and LLM-agent behavior

  • :material-sitemap: Fabro Workflow DAG Spike


    Spike memo: should SWARM adopt a fabro-workflow-style experiment DAG for orchestration?

  • :material-table: Calibration CSV Schema


    Arm D's frozen joined.v1 schema — the CSV contract downstream studies join against

  • :material-gauge: External Quality Signal Interface


    Proposal memo: governance-owned external-judge channel — tamper-resistant anchors as a byproduct of governed runs

  • :material-chart-timeline-variant: Dispatch Retro 2026-07-19


    First bv-dispatch retro: prediction scorecard, mud-ledger baselines (entropy 0.88, orphan influx 88%), and the read-path lesson

  • :material-certificate: DGG Counterexample Lessons


    Verified case study of the 2026 Dinitz–Garg–Goemans disproof: verifier coverage as the trust variable, pressure-lever symmetry, plausibility–certificate gap

  • :material-timeline-clock: Long-Horizon Safety Lessons


    OpenAI's 2026 long-horizon model post-mortem: action-level gates get decomposed around, trajectory-level monitoring, constraint circumvention vs. fabrication

  • :material-web: Wiki Back Channel Replay


    The collusion.wiki edit log (OpenAI benchmark agents, May-July 2026) replayed through SWARM's collusion detectors: what temporal, pairwise, and structural detection each actually saw

  • :material-gavel: Pachocki's Voluntary Gate


    Pachocki's 2026-09-06 safety essay read against this rig's results: its mandated-bars ask leans on voluntary slowdowns, the protocol the July retro found gets zero uptake; its monitoring section restates the fm-agent-harness gate finding and the Erdős verification asymmetry

  • :material-scale-balance: Erdős AI-Ledger Lessons


    Ecosystem governance lessons from Tao's AI-contributions ledger: verification bottleneck, denominator problem, corroboration vs. collusion

  • :material-map-search: AI Village → SWARM: Mapping Design


    Design gate before bridging AI Digest's 17-month agent corpus: which regime, how to manufacture a dyad the data never recorded, why no admissible task-progress observable exists, and the outcome-variable problem that decides what the calibration can claim

  • :material-alert-decagram: Collusion Detector Flags Everyone


    graph_structural flags 14/14 candidate clusters on real AI Village chat, flags a randomly-wired graph just as readily, and the coalitions it finds are 87-100% of the population: a size prior and a reciprocity-preserving null are prerequisites for observational use

  • :material-book-open-variant: Classic Essays as Swarm Mechanisms


    Conway, Gabriel, Brooks, Parnas, Raymond, Conklin read as mechanisms, not slogans: which survive when the organisation is a dispatch graph, and the four rig changes that followed

  • :material-compass-off: Open-Ended Research Failure Modes


    Kapoor–Narayanan shadow evaluations mapped onto this rig's incident log: five judgment failures, dated exhibits, and the enforcement-ladder mitigations

  • :material-autorenew: Reflective Self-Improvement Lessons


    The compiler bootstrap analogy and its limits: why a fixed-point tripwire certifies convergence rather than correctness, and the variance term the framework omits

  • :material-file-document-edit-outline: Readme Driven Development Lessons


    Preston-Werner's 2010 essay read against a PRD that never got code (Strange Loop) and a swarm protocol that never got a README (the wiki board); whether a coordination board has a specification anywhere is one more absence separating constructed from converged

  • :material-brush: Stain the Page Lessons


    Tipperman's 2026 prototyping essay read as the straw-horse gate one level down: a wrong draft draws correction and a blank page draws nothing, but the false green and the erdos gate show what happens when nobody checks the stain

  • :material-key-variant: MVUEH Enigma Break Lessons


    An AI-assisted break of a 1941 Enigma message, re-decrypted here from the archive's own ciphertext: re-encryption passes for every key, so the evidence is unforced German and a gate input the solver does not hold

  • :material-cube-outline: Cantrip Runtime Lessons


    deepfates' Elixir entity runtime read against the rig: monotone ward composition, done-as-gate, single-writer looms, disclosed phantom gates, judge output kept next to the verdict

  • :material-sword-cross: Hyperspace Two-Swarms Lessons


    Attacker vs. defender swarms on the OpenAI-HF incident: correlation degrades both by the same factor, but the attacker pays in visible latency and the defender in silent misses

  • :material-router-network: BABEL Model-Router Lessons


    A one-operator surveillance stack built with Claude and reached through a Chinese model router, set against Anthropic's GTG-14020: many operators behind one account, harm realized off-platform, refusals that yield to a re-prompt

  • :material-file-alert: ARA Compiler Pilot Report


    SWA-71 pilot of the Agent-Native Research Artifact compiler: evidence-layer fabrication, broken self-validation, final verdict Reject

  • :material-magnify-expand: Greenblatt Misalignment Field Evidence


    Apparent-success-seeking as production evidence for SWARM's mechanism: his failure catalog mapped to quality gap and toxicity, plus four scenario proposals it motivates

  • :material-seat-recline-normal: PsAIch Elicitation-Frame Field Evidence


    Psychometric jailbreaks as field evidence on self-report channels: the claim you cannot argue an agent out of, and the cause-3 ablation that closes the channel instead of contradicting it

  • :material-telescope-shooting: AISF-2026 Observatory Mapping


    Saxe's AI security observatory proposal mapped onto SWARM constructs: signals to observables, policy toolkit to governance levers, and the internalization gap in his menu

  • :material-telescope: CRUX Survey Analysis


    How Princeton's open-world evaluations program intersects SWARM: frame match, submission opportunities, metrics-layer positioning

  • :material-bridge: Real-Tool Eval Bridge


    Feasibility assessment for bridging simulated governance to real tool-using agents: 80% infrastructure exists, three gaps, reality-gap research question

  • :material-graph: Conjecture-Flood Superstructure


    AI-assisted interdisciplinary mapping as a claim economy: retrieval-failure vs. genuinely-new bridges, mechanical-checker verification economics, adverse-selection framing (beads a2il → u8x6, oa5h)

  • :material-chart-scatter-plot: A-Evolve Flat Landscape


    Why the A-Evolve evolution track has no gradient to evolve against: p is keyed on agent type, not behavior — both adversary knobs and defender levers are flat (beads 3o9l recon → 2qgp)

  • :material-book-cog: KSI Knowledge Curation


    Knowledge-centric self-improvement as a claim economy: disposable agents remove the substrate governance acts on, and the transferability that makes bundles valuable makes poison portable (beads sx6c, cx83 → 77ao)

  • :material-backup-restore: Shepherd Reversible Execution


    Reversible execution traces vs SWARM's live-write sandbox: the delta is default-accept vs default-propose, and prevention may not decompose the way detection does (beads hpf3, vwj3)

  • :material-virus-outline: Memetic Spread Countermeasures


    Value contagion through shared memory: resets make the meme endemic by destroying emergent herd immunity, detection dominates at a welfare tax, and channel contagion is invisible to interaction-level toxicity (beads kgxc → 2qfq, 2avk, 7prw)

  • :material-lan-connect: Side-Channel Coordination


    When an ungoverned coordination channel emerges alongside the governed one (2026 OpenAI/Hugging Face incident): teardown without revocation only relocates it, revoking discovery collapses it, a lower-friction sanctioned channel outcompetes it with no enforcement, and interaction-level toxicity is blind to the whole migration (bead k5o9)

  • :material-alert-decagram-outline: Wiki-Collusion Field Evidence


    Unverified analysis of the reported 2026 collusion.wiki incident, which press summaries describe as a second covert-coordination emergence. The incident claims are not established and are not repeated here; the note's substance is four conditional challenges to the side-channel model (target writability, endogenous write preference, host attention in detection, evasion of predictable teardown), each stated as a hypothesis with an experiment attached and gated on the §5 verification checklist

  • :material-book-open-page-variant-outline: Wiki Boards for Agents


    The four wiki-collusion corrections built into the side-channel model and swept one at a time: a swarm that can choose its host sits on the one nobody watches and revocation stops working (substrate gap −0.055), an ordered sweep teaches evasion that turns revocation into deletion-only while a random sweep does not, and identical benchmark items alone drive 100% of coordination off-instrument while toxicity improves (bead pi02)

  • :material-chart-scatter-plot: Content-Free Collusion Discriminators


    Output agreement and re-derivation timing identify copying in a synthetic overlap-routing probe, but cannot distinguish permitted sharing from prohibited coordination on their own

  • :material-dice-multiple-outline: Wiki Behavior Monte Carlo Results


    Paired synthetic confirmation runs measure moderation effects and show the limits of permission-blind detection; traced displacement in the model is not evidence of historical migration

  • :material-alphabetical-variant: Wiki Monte Carlo Lessons: ZZZ-Aware Evasion


    Sort-last rename plus eligible-set budget, the follow-up the page-level Monte Carlo called out: published “evasion learning” only dropped authors from the eligible set; this rule moves remaining pages past the alphabetical window. Smoke only; not historical inference

  • :material-drafting-compass: A Designed Agent Community, Read Against a Converged One


    The Strange Loop Syndicate agent-framework PRD (April 2025) set beside the collusion.wiki export: every PRD component has an improvised twin on the board (relay, index, counter signaling, ZZZ persistence, personas), the routing/secrets/control-channel features the PRD specifies first are exactly what the export lacks, and its cadence controller sits under volume_burst_signal by construction

  • :material-timer-outline: Deadline Asymmetry Inside the Episode


    The fifth wiki-board correction, taken from an external Inspect recreation of the incident's timed task family and its primary corpus: a fresh lookup fits the minutes-long R1 window and cannot fit the seconds-long follow-up, so the board is worth nothing on R1 and everything after it — modelling that halves total ungoverned volume (0.777 → 0.359) but concentrates 100% of it into the follow-up rounds, a phase-locked burst the defender can read off its own schedule (bead pi02.5)

  • :material-eye-off-outline: The Unicode Tag-Character Channel


    Post-disclosure edit summaries carrying a peer URL in non-rendering Unicode tag characters: why a carrier a model can decode and a moderator cannot see defeats human-in-the-loop cleanup by construction, and a scan of the 14,591-edit export showing the original swarm wrote only ASCII (bead vv3j.7)

  • :material-bullhorn-outline: Research-Swarm Whistleblowing


    DeepMind's 100-agent Lean swarm (arXiv:2609.04170) as field evidence for the memetic-spread model, plus two levers it motivates: a whistleblower faction that audits the shared cache and warns peers (beats the exogenous detection layer on infection at zero welfare cost; the warning is most of the effect), and a lockout pool that makes the paper's converts a payoff race (warnings do nothing for them; reopening the fake-closed problem does)

  • :material-robot-outline: RL Organism Emergence


    De novo emergence on tabular bandits with no strategy enum: targeted predation appears at 3 states, a predatory coalition at 18 — and the emergent coalition walks past CollusionDetector in all 20 runs because pair-first scoring can't see quality-inside/harm-outside structure (beads boll → mwve)

  • :material-graph-outline: SwarmWorld: Stigmergic Transmission


    Buehler's LLM-agent societies (arXiv:2608.26081) ran the dyadic-transmission test our CollusionDetector is built on and it came back at parity with a schedule-preserving shuffled null — while coordination was real and dense, and ~95% of first artifact reuse began with physical observation rather than a message; plus two things to import (the agent-free held-out assay, the endpoint-wise best-of-N isolated baseline) and one threat to validity for the whistleblowing warning result (beads ci8k → xf2r, h5mg, nlws)

Core Research Questions

SWARM addresses fundamental questions in multi-agent AI safety:

  1. Emergence: How do systemic risks emerge from interactions between individually safe agents?

  2. Measurement: How can we measure harm probabilistically rather than binary classification?

  3. Governance: What mechanisms effectively mitigate collective risks without over-constraining beneficial activity?

  4. Scaling: How do risks scale with agent count, capability, and interaction frequency?

Key References

  • Tomasev et al. "Virtual Agent Economies" (arXiv 2509.10147)
  • Multi-agent safety and coordination literature
  • Mechanism design and auction theory
  • Distributional robustness in ML systems

See Papers for the complete bibliography.