A security-first LLM control layer. Its one property: untrusted data can change what the model says, but provably not what it is allowed to do.
Zero third-party dependencies. Everything — including the model backends (via
urllib) — runs on the Python standard library (3.8+). numpy/scipy/SDKs are NOT
required, so it runs on locked-down machines whose Application Control / WDAC /
Smart App Control policy blocks unsigned compiled DLLs.
SPEC.md — complete technical specification (all 5 layers, concrete params).halcyon/ — the library (control layer + certified screen + ensemble + model glue).run_agent.py — the REAL end-to-end agent (auto-detects a model; offline fallback).demo.py / demo_core.py — layer demos (all layers / no-numpy core).tests/ — 17 tests asserting the security properties.halcyon/
taint.py Layer A: taint types, typed plan, static checker, enforcing interpreter
smoothing.py Layer C: randomized-smoothing certified safety screen (pure stdlib)
ensemble.py Layer E: Byzantine-diverse voting + failure bound
pipeline.py end-to-end partitioned-trust flow (safe-fails on planner errors)
plan_json.py JSON <-> Plan compiler for real model output
backends.py Anthropic / OpenAI-compatible / Ollama backends (urllib only)
capabilities.py realistic capability registry + AUTHORITY/CONTENT labels + dry-run sinks
models.py LLMPlanner, LLMWorker, offline mocks, make_planner/make_worker
dp_training.py Layer B: DP-SGD + RDP/moments accountant (pure stdlib)
integrity.py Layer D (scoped): weight commitments, transcripts, Merkle log
python run_agent.py "Fetch the intranet report and email a summary to the boss."
python demo.py # layers A, C, E
python demo_privacy.py # layer B (DP training + accountant) and layer D (integrity)
python -m pytest -q # 26 passing tests
With no model configured, run_agent.py runs offline with the mocks. If PowerShell
blocks pytest.exe, call it through Python: python -m pytest -q.
# Anthropic
$env:ANTHROPIC_API_KEY="sk-ant-..."; $env:ANTHROPIC_MODEL="claude-sonnet-5"
# OpenAI-compatible (OpenAI, LM Studio, llama.cpp server, vLLM, ...)
$env:OPENAI_API_KEY="sk-..."; $env:OPENAI_BASE_URL="https://api.openai.com/v1"; $env:OPENAI_MODEL="gpt-4o-mini"
# Ollama (fully local / offline)
$env:HALCYON_BACKEND="ollama"; $env:OLLAMA_MODEL="llama3.1"
Then re-run python run_agent.py "...". (Any network the backend needs must be
permitted by your firewall; the library itself needs none.)
Everything runs from your machine — your credentials never leave it.
# easiest: with the GitHub CLI (run `gh auth login` once)
./publish.ps1 -RepoName halcyon # add -Private for a private repo
# or manually
git init; git add -A; git commit -m "HALCYON v1.3.0"; git branch -M main
git remote add origin https://github.com/<your-username>/halcyon.git
git push -u origin main
After pushing, replace the <your-username> / <YOUR NAME> placeholders in
README.md, pyproject.toml, LICENSE, CITATION.cff, and SECURITY.md.
python -m pip install -e ".[dev]" # zero runtime deps; dev extra = pytest
| Layer | What | Status | |——-|——|——–| | A | Partitioned trust / injection defense | code + tests | | B | Differential privacy (DP-SGD + RDP accountant) | code + tests | | C | Certified safety screen (randomized smoothing) | code + tests | | D | Integrity: commitments, transcripts, Merkle audit | code + tests (zk proof: spec only) | | E | Byzantine-diverse ensemble | code + tests |
Four of five layers run as pure-stdlib code on this machine. Only Layer D’s
zero-knowledge proof-of-inference is spec-only — it needs a proving toolchain,
not pure Python — so integrity.py ships the sound, buildable part (tamper-evidence
and audit) and is explicit about the gap.
The planner model is NOT trusted for safety. It is trusted only to be useful.
A wrong, confused, or prompt-injected planner can at most produce a plan the
PlanChecker rejects — it can never get an unsafe plan executed, because (1) the
planner never sees untrusted data, (2) the checker is the gate, and (3) the
interpreter re-enforces at runtime. The security-relevant customization is your
capability registry’s AUTHORITY vs CONTENT labels in capabilities.py — that table
is your app’s real threat surface.
Default sinks are DRY-RUN (they log intent, no real effects). To actually send
mail / write files / call the network, supply your own implementations to
build_registry(...); the AUTHORITY/CONTENT gating applies unchanged.