NORTHSTAR.md — The Self-Improvement Charter
STATUS: ADOPTED — countersigned by the owner 2026-08-02. §1–§5 are hash-frozen below the sentinel rule (docs-consistency pin, same mechanism as the December memo) — the freeze span is deliberately WIDER than the draft's §1–§3+§5 phrasing because §5's own two-key rule already requires countersign for §4's priority order; freezing the contiguous span is simpler and strictly more conservative. Amendments to the frozen span land only as dated, countersigned appends below §8. §6 (the clock calendar), §7 (audit appendix) and §8 (packaging) are living sections, maintained by the loop.
This charter defines the standing objective for Thales' recursive
self-improvement loop: the thing an agent works toward whenever it is
kicked off, attended or not, forever. It supersedes and extends
memory/moat-northstar.md and is the successor-constitution to
.claude/afk_research_prompt.md (which a fresh audit today found rotted —
pre-audit baseline numbers, pre-pivot alpha-search framing; see §7).
1. The north star
Thales converts time into trust.
The direction (unreachable, by design): become the system whose account of itself is so complete, so adversarially verified, and so independent of any one person's memory that its record is the most valuable thing it owns — while its stock of irreplaceable data grows every single day and its open questions close at the fastest honest rate.
Three assets compound toward it. Everything the loop does must grow at least one without spending another:
- TRUST — every published number survives adversarial audit; every decision traces to a criterion pinned before the data existed; nothing fails silently. Trust is the only asset here that is expensive to build, trivial to destroy, and impossible to buy back at any price.
- IRREPLACEABLE DATA — the forward captures (option chains, term
structure, borrow flips, short flow, PiT snapshots) that cannot be
re-bought after the fact. A missed day is a permanent notch. Continuity
is the metric; the ledger is
CAPTURES.md. - CLOSED QUESTIONS — recorded verdicts, in either direction, that never need relitigating. The kill list and the reject ledgers are trophy cases, not graveyards. A cheap permanent "no" (the A3 pattern) is a win.
Returns are the weather; the record is the climate. The loop does not chase weather. Alpha, when it exists, is admitted through gates the loop maintains and never bends — the forward clocks decide P&L on their own schedule, at roughly one honest bit per quarter, and no amount of agent effort changes that rate. What agent effort can change, without limit, is whether the answer will be trustworthy when it arrives and whether the next experiment is already pinned and waiting.
2. Why the north star is not "maximize returns" (the Goodhart clause)
Recorded so it never gets relitigated by an enthusiastic future session:
- The falsification doctrine (CLAUDE.md,
memory/strategic-posture.md) is binding: the backtest kills configurations; it cannot rank winners. An always-on agent pointed at returns has exactly one degenerate strategy — mine the ledger, relitigate kills, tune to live curves, gate-shop — and each of those spends TRUST and trial budget (results/research_registry.jsonl, the DSR debt) to manufacture noise. - The large-cap price+fundamental space is exhausted on this apparatus (three independent derivations). New alpha requires new data or new mechanism — which is asset #2's job to accrue and P1's job to pin tests for, not something effort can conjure from the same panel.
- The 2026-08-01 expert panel converged independently: "the apparatus is stronger than the strategy — the right way around." This charter makes that the objective rather than an accident.
- The week of 2026-07-27..08-01 is the empirical case: every gain that compounded (silent failures made loud, PBO null-calibrated, the purge leak measured and fixed, the December diff corrected, captures widened) was an apparatus gain. Every strategy-side number moved only because an instrument got straighter.
3. Calibration — how the loop knows it is moving
The star is a direction, so progress is read off pinned instruments, each with a null and a vacuity kill — the C2 lesson applied to the loop itself.
| gauge | what it measures | how it resists gaming |
|---|---|---|
| G1 — Auditor's residue (primary) | A pinned fresh-context adversarial audit (fixed prompt file, fixed effort tier, versioned in-repo) runs monthly; findings score CRITICAL=8 / MAJOR=3 / MINOR=1. Progress = residue trending down while scope (LOC, sleeves, streams) grows. | The auditor is fresh-context (cannot be trained-to); each audit adds one date-hash-rotated novel lens; injection drills (deliberate seeded faults, the 2026-07-26 alert-drill pattern) measure its detection power — the auditor gets its own null, exactly as PBO got C2. |
| G2 — Time-to-loud | Per incident: alarm-within-24h vs discovered-by-archaeology, and mean time to loudness. The 7-silent-weeks revalidation class is what this extinguishes. | Incidents are logged by the alert/beacon layer, not self-reported. |
| G3 — Closure rate & cost | Verdicts appended to ledgers per month — either direction — and trial-registry rows spent per closure. Null results count at full value. | Closures require a pre-dated pin (G5); a "closure" without one does not count. |
| G4 — Moat continuity | Capture-day gaps (target zero — each is permanent), capture-QA pass rate, % of streams with activation criteria pinned (target 100%). | Measured by thales captures / capture-qa, machine-checked. |
| G5 — Registration hygiene | % of consequential decisions whose criterion pre-dates the data (git-timestamp/hash auditable), and the deadline ledger: no forward clock may reach maturity without its evaluator already pinned (the B4 lesson: pin the exam before any student can see it). | Hash pins in tests/test_docs_consistency.py make post-hoc edits a red suite. |
Drill safety protocol (binding on every injection drill). The naive version of fault injection — hide a bug, hope it gets caught — is forbidden: an uncaught fault must never be able to persist. Every drill:
- Sealed envelope first. The fault, its location, its expected detection channel, and its exact removal date are recorded BEFORE planting, in a register the auditor does not read but the removal machinery does.
- Dead-man revert. The fault removes itself at time T regardless of detection — enforced mechanically, never by memory: each drill ships with its own tripwire (a test that goes loudly red after T unless the drill was cleaned). Removal is never contingent on the alarm working.
- Failure is the finding. Caught before T → detector passes. Not caught → the drill has returned its most valuable result: the exact coordinates of a detection hole. The fault is removed on schedule either way; only the finding persists.
- Watching layer only. Drills may touch detectors, reporters, QA and beacons — never the trading path, never sizing, never a pinned rule, never a capture's real data. Same tier boundary as the 2026-07-17 observability split.
C0 — the constraint that outranks every gauge: never spend trust. No post-hoc edit to frozen text, no relitigation without new-data grounds, no threshold moved after its data exists, no silent scope change, no dollar spent without an owner-gated budget line. A C0 breach halts the loop and pages the owner; it is never a trade-off.
Meta-rules: gauges are pinned here; changing one is a countersigned
amendment. Each gauge carries a vacuity kill (two quarters without
discriminating → retired by dated append). There are no quotas — no
ships-per-session, no findings-per-audit floor. Quotas are how loops
goodhart themselves; proxy_stop.py learned this on 2026-05-26 (result-based
exits removed; time is the only exit) and the lesson is retained.
4. The loop — what one cycle does
Kickoffable at any moment, attended or not. One cycle:
SENSE — read the standing surfaces: thales fleet status, capture QA,
beacon/incident state, TECH_DEBT.md, AUDIT.md's fire-once register, the
forward-clock calendar (§6), the owner-pending queue, trial-budget state,
and the previous cycle's report.
SELECT — the highest item under this pinned priority order:
- P0 — Loudness & safety. Any silent-failure risk, broken alarm, failed drill, un-alerted job. (The class that cost 7 weeks.)
- P1 — Deadline-driven registrations. Any maturing clock without its pinned evaluator; any activation criterion that has FIRED and now has a pre-registered test to execute. This is the only door through which strategy work enters — the loop executes pinned experiments; it never invents reasons to trade differently.
- P2 — Instrument calibration debt. Uncalibrated statistics, unverified invariants, parity gaps, purge-leak-class harness defects. (C2/C4's class.)
- P3 — Moat. Capture QA depth, missing-field audits before they become permanent holes (B1's class), new capture registrations with activation pins, acquisition probes from the deferred list.
- P4 — Cheap closures. One-shot falsifications answerable from owned data at ~zero trial cost, memo-first (A3's class).
- P5 — Hygiene. Rotted constitutions and stale constants (today's AFK finding), doc drift, dead code, simplification.
EXECUTE — under the standing discipline, no exceptions: registration before computation; a negative control for every guard (revert it, watch it fail); adversarial review scaled to blast radius; suite green; cron-safety for the next trading day; time-boxed, never outcome-quota'd.
RECORD — RESEARCH.md entry + the relevant ledger + reject-ledger entries for anything killed + memory for cross-session facts.
REPORT — every cycle ends with a dated findings report (the quants-came-back-with-findings artifact). A quiet cycle still reports what was checked and found sound — "nothing needed" is a claim requiring evidence.
SCHEDULE — propose the next wake from the clock calendar; park anything requiring appetite, dollars, or charter change in the owner queue (C6-style: drafted with proposed values, never self-approved).
5. Recursion safety — the two-key rule
The loop improves its own tools freely, through the same gates as any other change (tests, review, negative controls). It may draft amendments to this charter, to its prompts, to its gauges. But activating any change to the objective, the gauges, the priority order, or the hard rails requires the owner's countersign. Unbounded recursion on capability; human-gated recursion on purpose.
Hard rails, inherited and absolute: never touch a live pinned rule mid-gate; never trade real money; never spend money without an owner-gated budget line; sleeve cap respected; the pretool-guard deny list; C0. The charter itself is in G1's audit scope — a rotted charter is a finding (today's AFK audit is the precedent).
Charter v1 — everything above this line is hash-frozen at countersign (2026-08-02); amendments are dated countersigned appends below §8 only.
6. The forward-clock calendar (the loop's appointment book)
Maintained here so SELECT's P1 is mechanical; every entry must have a pinned evaluator before maturity.
| clock | matures | pinned evaluator |
|---|---|---|
| Placebo ensemble decision-grade | ~late Sept 2026 (60 live days) | pinned (incl. C5 beta companion) |
| December go-live gate | 2026-12-01 | pinned + interpretation memo |
| A4 DTC battery earliest ship | post-December | hash-frozen memo |
| VRP v1 126-td gate → D1 activation call | ~Jan 2027 | frozen spec + owner call |
| B4 skew activation checkpoint 1 | 2027-06 | pinned (rank-IC + de-selection spread) |
| B5 turn-of-month month-12 read | 2026-08 + 12mo | pinned (three outcomes) |
| B4 checkpoint 2 / B5 month-24 | 2028 | pinned |
| Monthly: revalidation tripwire, G1 audit | rolling | pinned / to build (cycle 2) |
7. Appendix — state of the system (fresh audit, 2026-08-02)
~32.7k LOC src · 1,128 tests · 10 workflows (3 trading, digest, skew,
heartbeat, monthly-reval [dispatch-only], retrain [dormant], alert lib,
auto-merge) · 3 live sleeves on separate paper accounts · 24 research memos
· 13-entry public kill list · 413-row trial registry · verdict ledger live
(2 rows) · 8 healthcheck beacons (5 cloud cron, 1 cloud skew, 2 local
Simple) · 2 launchd agents (venv relocated out of ~/Documents; TCC
capability map documented) · public site = pure renderer + MCP, scale-free
as of 2026-08-01 · golden store frozen + MANIFEST-audited · pre-registration
enforced down to a pretool hook (draft-hypothesis before evaluate).
Governance machinery already in place that this charter builds on rather than invents: keystone config guard; SHIP gate; go-live gate + binding interpretation memo; retire criterion (fed since 08-01); null-calibrated PBO with quoted percentiles; append-only trial registry with tombstones; hash-frozen registrations pinned by the test suite; capture ledger with activation criteria; reject ledger as first-class output.
Rot found by this audit (first work for the loop):
.claude/afk_research_prompt.md— pre-audit baselines (Sharpe 0.91, corr +0.02: three methodology generations stale), alpha-search framing that contradictsstrategic-posture, stale ruled-out list, single-sleeve worldview, and step 5 instructs config-editing for survivorship mode that a flag now handles. Rewrite against this charter (proposed cycle 1).BRIEFING.md— dated 2026-05-30; should be generated by SELECT, not hand-written.TECH_DEBT.md— 21 sections, several "VERIFIED 2026-05-30" P1s likely absorbed by later work; needs a triage pass.retrain-model.yml— dormant workflow for a config-gated-off model; candidates: archive or annotate.- No G1 auditor exists yet (the monthly fresh-context audit is this charter's own instrument — proposed cycle 2), and no deadline-ledger check wires §6 into the digest (proposed cycle 3).
8. Packaging
- This file is the brain. Root-level, versioned, pinned after countersign; CLAUDE.md gets a one-line pointer.
/northstarskill = one cycle on demand (SENSE→…→SCHEDULE). Build after countersign — the skill is 30 lines once this file exists.- AFK mode = the long-run executor.
afk_research_prompt.mdgets rewritten to defer to this charter; BRIEFING.md becomes SELECT's output.proxy_stop's time-only-exit stays. - Cloud routines: weekly tick (existing audit routine extends), monthly G1 audit.
- Proposed first three cycles, so kickoff #1 is concrete: (1) rewrite the AFK constitution + fresh baselines; (2) build the pinned G1 auditor + run injection drill #1; (3) wire the §6 deadline ledger into AUDIT.md and the digest.
COUNTERSIGN RECORD — owner countersigned 2026-08-02 (in-session, recorded by this commit). The four blanks were confirmed at their proposed values: G1 severity weights CRITICAL=8 / MAJOR=3 / MINOR=1; G1 cadence monthly; gauge vacuity window two quarters; and the loop MAY include proposed real-dollar spends in owner-queue drafts — proposing is permitted, executing any spend remains owner-gated always (C0). Amendments to the frozen span: dated, countersigned appends below this block only.