The finalized chain loop
The production loop is a non-emitting, restart-safe intake and qualification controller. It does not accept a shell evaluator, keep a JSON scoring ledger, or submit weights after each pass.
One pass
run_pass(...) performs these operations in order:
- Bind chain scope. Open the SQLite store against the chain genesis hash and netuid. A database created for another scope is not reused.
- Read finalized history. Continue after the durable cursor and reconstruct exact reveal priority from finalized storage and canonical event positions.
- Reserve before transport. Persist every arrival in chain order before any fetch. Slow hosting therefore cannot rewrite priority.
- Fetch privately. Accept HTTPS only, validate DNS and every redirect, enforce archive limits, extract regular files safely, and rederive the committed hash.
- Classify and fingerprint. Parse the target-scoped proposal; a bundle the component parser rejects is refused. Copy identity covers the submitted delta, never the validator's incumbent stack.
- Publish immutably. Copy the validated private tree to a content-addressed worker publication and reopen it before use.
- Reconcile copies. Compare durable fingerprints in finalized order. This step is separate and idempotent so a crash between publication and copy disposition cannot bypass priority.
- Screen and qualify. If a registered arena service was injected, run its staged screens, use the routing-only resident lane where applicable, form a capacity-bounded cohort, execute authoritative resident qualification, and persist outcomes.
- Settle retained pairs. Lease economically unblocked, independently reproduced candidates and apply the resulting settlement plan transactionally.
The pass returns counts and dispositions. It never opens a wallet or calls
set_weights.
Reservation state machine
The row status is an operational control signal, not merely a progress label:
qualified means two matching PASS qualifications have been retained. The associated
settlement candidate then has its own transactional state (pending, leased, and a
terminal economic disposition such as crowned, held, neutralized, or
discovery_bounty). A reservation can remain qualified while settlement decides its
economic outcome; do not infer a crown from the reservation status alone.
Several terminal paths are omitted from the diagram for readability: an operator may
explicitly expire sufficiently old inactive work, copy reconciliation can turn a later
submission into failed, and bounded retry exhaustion leads to held.
Public CLI: intake only
The supported standalone command is:
cacheon chain-validate \
--netuid 307 \
--network "wss://test.chain.opentensor.ai:443" \
--intake-db chain_intake/intake.sqlite3 \
--private-root chain_intake/private \
--publication-root chain_intake/worker \
--audit-log chain_intake/chain-audit.jsonl \
--intake-only \
--onceRemove --once to run continuously; --interval controls the delay between passes.
chain-validate accepts only its declared intake and arena schema. Chain-signing
credentials, external evaluator commands, scoring policy, and weight publication belong
to separate authorities and must not be added to the validator-loop service.
Without --intake-only, the CLI rejects startup unless its Python caller injects an
exact ArenaServiceRegistry and selects a registered --arena-id:
from cacheon.chain.validator_loop import run_validator
run_validator(
subtensor,
netuid,
intake_db="chain_intake/intake.sqlite3",
private_root="chain_intake/private",
publication_root="chain_intake/worker",
audit_log="chain_intake/chain-audit.jsonl",
arena_registry=registry, # constructed by reviewed deployment code
arena_id="production-arena-id",
intake_only=False,
)This is an integration boundary, not a copy-paste complete deployment: the repository
does not provide the production provider represented by registry.
In daemon mode, run_validator contains pass-level validator faults. It logs the full
exception, increases the sleep multiplier up to six times the configured interval, and
resets the failure count after a successful pass. The Python API default stops after ten
consecutive failures so a supervisor can intervene. Candidate dispositions already
committed before the process-level exception remain in SQLite; the loop does not roll
the entire pass back as one transaction.
HTTPS intake boundary
The on-chain payload is a canonical JSON object containing schema version, lowercase SHA-256 content hash, and an HTTPS URL. It is limited to 1024 UTF-8 bytes.
Production fetch enforces:
- HTTPS and TLS 1.2 or newer;
- globally routable resolved addresses;
- connection to a reviewed address while retaining TLS SNI and hostname checks;
- validation of every redirect, with at most five redirects;
- 64 MiB downloaded archive, 256 MiB extracted content, and 4096 members;
- 16 MiB per regular file, 8 MiB per inspectable source file, and 32 MiB aggregate inspectable content;
- raw gzip/tar preflight before
tarfilematerializes metadata, including 64 KiB per PAX/GNU extension header and 1 MiB aggregate extension payload; - no symlinks, hardlinks, special files, duplicate/path-conflicting members, or path traversal; and
- a bounded transfer deadline.
file:// exists only behind explicit test helpers. It is not accepted by
chain-submit, payload decoding, or the production fetch function. Plain HTTP is not
accepted at all.
Private and worker storage
The fetch root is validator-owned private storage. The code requires owner-private directories and files and never mounts this mutable intake tree directly into a worker.
Publication creates a separate carrier:
- all bytes are proven to participate in the committed identity;
- files are sealed read-only and directories are non-writable;
- the destination address is content-derived; and
- reopening independently rederives the publication and content hash.
Treat both roots as operational data, not as interchangeable caches.
Redacted journal and private recovery snapshots
When --audit-log is configured (the CLI default), each completed pass appends and
fsyncs one canonical JSONL record. The record contains finalized position,
content-derived reservation and receipt digests, counts, and bounded disposition
classes. It deliberately omits URLs, hotkeys, candidate bytes, exception messages,
wallets, credentials, and ambient environment. A fault record retains only the
exception type and consecutive-failure count.
SQLite remains the state machine. The journal is supplementary operational chronology: an audit append failure is reported but cannot cause the loop to replay a pass whose SQLite transitions already committed.
Operator reservation diagnostics
The CPU validator can answer a miner's status question without stopping intake or opening a second writable controller:
cacheon chain-reservation-status \
--intake-db /srv/cacheon/state/intake.sqlite3 \
--audit-log /srv/cacheon/state/chain-audit.jsonl \
--reservation-id <64-HEX-RESERVATION-ID>Use --content-hash or --miner-hotkey only when it identifies exactly one retained
row. Ambiguous selectors are refused. --json emits the same privacy-safe record for a
support tool. Both output forms omit proposal URLs, private filesystem roots, and raw
exception messages.
Read the result in this order:
arrival_authorityis the finalized-chain ordering fact. The order key is block, event index, event subindex, hotkey, and content hash.queueis present only for queued or active work. A numeric position ranks actual selectable work: reproduction first, then primary, with finalized order inside each class. An active screen or qualification has no queue position. A durable remote lease isleased, has no queue position, and is excluded from the remaining waiting depth;evaluation_leasesupplies its stage, generation, cohort position, and expiry block without exposing the private worker owner. Promoted qualification work can be an indivisible retry group or bounded cohort, so it is not assigned a misleading single-row rank.- A screen
rejectis explained by the terminal typed stage, grade, and evidence digest inscreens. A qualification outcome is explained by its persistedPASS/FAIL/NO_DECISION, reason code, report or failure digest, and typed attempt reference when one was retained. - Attribution is derived only from a persisted typed decision. A row whose status is
failedbut which has no persistedFAILremainsunattributed; status text alone is never used to blame candidate code. evidence_limitationsis mandatory support context. In particular,qualification_failure_retained_by_digest_onlymeans the current schema retained an infrastructure failure product's digest but not a reopenable artifact reference. Report it as validatorNO_DECISION, quote the digest, and do not guess at a more specific cause.
The last case is a real current limitation: this command cannot reconstruct bytes that were never durably referenced. Private evaluator logs must therefore remain under the deployment's log-retention policy until typed failure-artifact retention and archive indexing cover every infrastructure path. The command makes that gap visible; it does not claim the gap is closed.
For support across every retained submission by one miner, use
chain-miner-report --miner-hotkey <HOTKEY>. It composes the same privacy-safe
reservation facts with duplicate-replay history and prints the stated cause and
next action for each row. It does not infer candidate blame from status text and
cannot recover failure bytes that were retained only by digest. See the
CLI reference.
The private recovery mirror is also outside the live transaction path:
cacheon chain-snapshot \
--intake-db chain_intake/intake.sqlite3 \
--audit-log chain_intake/chain-audit.jsonl \
--object-store-bucket <PRIVATE_BUCKET> \
--object-store-endpoint <S3_COMPATIBLE_ENDPOINT> \
--sealed-input qualification-inputs=/srv/cacheon/sealed-inputs
cacheon chain-snapshot-verify \
--manifest-key <PRINTED_MANIFEST_KEY> \
--object-store-bucket <PRIVATE_BUCKET> \
--object-store-endpoint <S3_COMPATIBLE_ENDPOINT>chain-snapshot takes a consistent online SQLite backup, checks integrity and foreign
keys, discovers database-referenced immutable worker publications and retained
settlement qualification artifacts, and adds only explicitly named sealed inputs.
Every object is digest-addressed, bounded during download, and reopened before the
manifest is accepted. Models, OCI images, wallets, credentials, caches, unredacted
logs, and unrelated evidence roots are not auto-discovered.
chain-snapshot-verify restores into a fresh private staging root (temporary by
default), rechecks SQLite, publication receipts and hashes, evidence references, and
the redacted journal, and emits a restore map when a retained --restore-root is
requested. It never overlays live state. Schedule snapshots separately from validator
passes so remote object-store failure cannot enter the SQLite controller's commit
path. Protect the archive with a private bucket or policy-isolated private prefix and
exercise restore regularly.
Evaluation-lease ownership
FinalizedIntakeStore owns durable evaluation-lease state. FIFO selection order,
reproduction priority, qualification cohort formation, lease generations, heartbeat
compare-and-swap, expiry, and infrastructure release are store policy. No other
module implements a second lease state machine, and deployment tooling must not
mutate lease rows with raw SQL.
cacheon chain-evaluation-lease is the tracked one-shot operator adapter over that
API: preview, claim, heartbeat, and infrastructure release, with all
authority coming from one sealed owner-controlled config file. It is not an
evaluation worker, daemon, or scheduler; see the
CLI reference.
Site orchestration — launch wrappers, tmux composition, endpoints, wallets, exact filesystem paths, and sealed production configs — stays in the private deployment tree outside this repository. Tracked code owns lease semantics; private operations own only identities and launch composition.
Remote worker transport
Remote execution of leased work is a durable-spool transport, not a second
evaluation authority. chain/remote_worker_registration.py binds one worker
epoch — endpoint, pinned host keys, commissioned READY receipt, worker
readiness, physical lane, interpreter, and shared credential — under one
semantic digest. The immutable worker carrier places the validator's
.cacheon-native-artifact.json receipt beside the miner's committed source;
bundle_hash.committed_content_hash is the one canonical rehash of a
receipt-bearing carrier back to the chain-committed identity, excluding
exactly that top-level receipt and nothing else. chain/remote_worker_spool.py owns the sealed
request/result carriers and their verification;
chain/ssh_worker_transport.py shuttles them over host-key-pinned SSH and
implements the authenticated transport the remote evaluation dispatcher uses
for both screen and qualification; chain/remote_worker_pod_service.py
supervises one persistent pod adapter per epoch and parks the epoch on its
first command-level adapter failure rather than restarting into an unproven
resident model. Transport, pod, and adapter failures surface as
infrastructure no_decision records that release the durable lease without
consuming an evaluation attempt.
The tracked B300 adapter has two closed construction modes. Screen-only mode
executes through the commissioned screen deployment and refuses qualification
before resident work. Persistent --serve mode may additionally load one
digest-exact qualification-capabilities factory; it then constructs the shared
commissioned B300 service, resolves each authenticated promoted cohort, and
executes remote qualification through the same READY-bound worker and durable
continuation store. One-shot mode cannot commission qualification. Deployment
wrappers supply every installed path as an explicit argument; active endpoints,
credentials, sealed capability bytes, and process composition stay in the
private operations tree.
Standing CPU supervisor
python -m cacheon.chain.standing_cpu_supervisor --config <path> is the
standing CPU daemon over those pieces. Its sealed, closed, owner-controlled
config names the screen-dispatcher config (chain/mainnet_screen_dispatcher.py
supplies the config schema and the dispatcher builder) and the
recoverable-qualification authorities to compose, and optionally the settlement
network. The screen stage reopens the intake-only validator's durable finalized
cursor read-only — rejecting scope drift, regression, and hash changes — claims
exactly one durable screen lease at a time, and hands the typed request to the
authenticated spool transport. The required ArenaService provider slot is
filled by a digest-exact remote-only proxy whose execution methods always fail
closed. The qualification stage resumes the same durable request across
restarts rather than restarting the experiment. Stage faults tear down the
constructed authority and rebuild it under bounded exponential backoff; every
status change is one canonical-JSON line on stdout.
enable_settlement installs the transactional settlement stage. When enabled,
settlement_network must name the finalized-head endpoint used to clock lease
and commit; the stage refreshes that clock immediately before each stateful
boundary and never opens a wallet. A commission may stage a nonempty
settlement_network while the flag remains false so arming is a reviewed
one-field change; enabling settlement with an empty endpoint is refused.
enable_weights installs the eval-side weight-offer push stage
(chain/standing_weights_stage.py) and requires weights_stage_config, an
absolute path to a second sealed, closed, owner-controlled file with schema
cacheon-standing-weights-config-v1 and exactly these fields: network
(explicit wss:// finalized-head reader), fallback_endpoint (empty or
wss://), push_url (http(s) serve-weights offer endpoint),
push_credentials (owner-only path to the push credential set),
attribution_hotkey, half_life_blocks, discovery_lifetime_blocks,
discovery_pool_ppm, refresh_blocks, and burn_hotkey. Every
refresh_blocks the stage reads the finalized head and metagraph, reopens the
intake store, and pushes the current V1 offer: the real projection whenever an
active reward claim, a crowned arena, or an activated composition exists;
otherwise the full-pool burn offer to burn_hotkey when that field is set, or
the builder's crownless refusal as a stage error when it is empty. The stage
never signs; the serve-weights lane owns readback and the follow-weights
signer decides what reaches the chain. Naming weights_stage_config while
enable_weights is false is refused, as is the reverse.
Durable reservation states
The store makes work and failure class explicit:
| Class | States |
|---|---|
| Active | reserved, fetching, transport_retry, published, screening, promoted, qualifying, reproduction_pending |
| Terminal | failed, expired, qualified |
| Operator/retry disposition | held, no_decision |
Default intake policy bounds include queue size, per-hotkey and per-target admission, transport and qualification retries, cohort size, epoch cutoff, and expiry. These are code defaults, not a promise that they suit every deployment; an operator should review them alongside arena capacity.
The default IntakePolicy values are:
| Bound | Default | Effect |
|---|---|---|
| Epoch / cutoff | 360 / 30 blocks | Arrivals in the cutoff tail are admitted into the next epoch |
| Pending queue | 256 | New valid arrivals beyond the bound fail admission deterministically |
| Per hotkey / epoch | 16 | Limits one submitter's intake occupancy |
| Per target / epoch | 64 | Applied after the target is resolved from submitted bytes |
| Transport / qualification attempts | 3 / 3 | Exhaustion produces a retained hold rather than infinite work |
| Controller cohort | 8 | Bounds fetch, screening, and qualification selection per pass |
| Finalized-block expiry SLA | 500,000 blocks | Keeps queued work for roughly 69 days before automatic stale-state expiry and sets the minimum age for explicit expiry |
The expiry SLA is only meaningful against the queue's service rate. The former 10,000-block bound covered only about 51 reservations at the measured service time and expired 178 queued rows without verdicts. The 500,000-block default keeps automatic expiry as a last-resort stale-state bound rather than a normal capacity disposition. Treat any cohort expiring without being reached as a capacity fault, not a miner outcome.
An exact cohort that expired because the validator worker was unavailable can be
readmitted with chain-evaluation-lease requeue-expired --authority <SEALED_JSON>.
The closed authority binds the reservation IDs, retained-result IDs, and the
fixed validator_worker_unavailable reason. The store restores each row to its
durable published or promoted lane and grants a fresh finalized-block SLA
without deleting history. The bounded refresh budget fails closed; an
owner-escalated repeat must be explicit in a newly sealed authority. This is not
a generic expiry undo.
Arena capacity is an additional bound. Its queue age/depth, active-screen, active-qualification, cohort, and retry limits are content-bound in the service manifest. Changing either policy changes operational behavior and should be reviewed and recorded; the code defaults are not calibrated economics.
The controller applies the finalized-block SLA on every pass, including retained-only
passes, and inside intake and settlement transactions that depend on unresolved priority.
Eligible reserved, transport_retry, published, promoted,
reproduction_pending, held, and no_decision rows expire automatically when their
arrival or retained-progress block reaches the bound. In-flight fetching, screening,
and qualifying rows are not aged out underneath active work. A first retained PASS
records a fresh finalized progress block and starts a full bounded reproduction window
from that block. Legacy retained evidence with an unknown progress block, including the
dedicated schema-3 migration hold, remains fail closed for explicit operator disposition.
This prevents slow reproduction from losing its complete SLA while preventing one old
PASS from becoming a permanent priority veto.
Eval-cost admission policy
Eval-cost admission is deliberately separate from the shared IntakePolicy used by
screen and evaluation-lease services. chain-validate --eval-cost-tao-rao controls the
required transfer_keep_alive amount and defaults to 0 (off). Quote TTL defaults to
300 blocks and the payment-to-reveal window defaults to 7,200 blocks.
A v2 reveal may attach a payment pointer. When the gate is enabled, intake rebuilds the remark from that reveal's hotkey, content hash, and netuid; only that triple can spend the pointer. The paying coldkey is not the claimant. When the gate is disabled, the pointer is ignored for payment accounting: unverified coordinates are neither consumed nor allowed to pre-claim a future payment.
Byte-identical resubmissions replay their prior verdict before any lease is claimed:
a bundle whose exact content hash already reached a terminal FAIL under the exact
current arena service digest inherits that FAIL (reason
duplicate_of:<reservation>:<original reason>) and costs neither a screen nor a
qualification. A prior PASS is never replayed — settlement requires an independently
bound PASS pair, so a resubmitted winner queues for a real evaluation. Any changed
byte, or any change to the arena, produces a fresh evaluation.
An operator can grant one artificial make-good with
chain-eval-cost-credit. The oldest unspent credit for that hotkey admits one
otherwise-unpaid reveal and is consumed inside the same admission transaction.
Other admission failures leave it unspent, and a reveal that cites a payment
pointer must pass ordinary payment verification instead of falling back to a
credit. Credit grants are private validator mutations with an audit note, not a
miner-controlled payment token; see the
CLI reference.
Verdict and retry semantics
The controller maps failures according to where authority was lost:
| Point of failure | Stored disposition | Retry behavior |
|---|---|---|
| Invalid chain payload, unpaid or invalid eval-cost payment, unsafe archive, content-hash mismatch, malformed proposal | failed / FAIL | None; attributable intake failure |
| Eval-cost payment lookup RPC/decode blip | pass aborted; cursor unchanged | Retry the pass; do not fail the miner |
| Transient HTTPS/DNS or immutable-publication storage fault | transport_retry / NO_DECISION | Retry until the transport budget, then held |
Static/build/ABI/graph/serving screen FAIL | failed / FAIL | None under that screen authority |
| Screen timeout or inconclusive evidence | Retry in the same primary or reproduction lane | Arena screen budget decides retry versus hold |
| Qualification plan/runner/raw-speed failure affecting a registered cohort | NO_DECISION for every member plus a persisted bisection plan | Cohort halves are retried to isolate poisoning without assigning losses |
Per-candidate post-attempt NO_DECISION | Retained report plus one-candidate requeue | Retry in primary or reproduction lane |
First complete PASS | reproduction_pending; no settlement candidate yet | Fresh screen and qualification required |
Second matching complete PASS | qualified; paired candidate becomes settlement-pending | Settlement leases it when earlier economic blockers clear |
The qualification retry counter counts retained qualification dispositions. The screen counter counts retained screen attempts. Restarting the service does not reset either. Likewise, changing a reason string or moving files does not create a fresh economic identity.
Restart behavior
On restart, the store does not pretend interrupted work completed:
- interrupted fetch or qualification becomes
heldwithNO_DECISION; - an interrupted screen returns to the appropriate retry lane; and
- an expired settlement lease returns to pending with a new generation.
The finalized cursor, reservation identities, and immutable publications make repeated passes idempotent. Validator/storage faults should produce retry or hold, not a miner loss. A supervisor can restart the loop, but must not delete or hand-edit the database to “unstick” it.
Recovery is intentionally conservative:
fetchingandqualifyingbecomeheldwithNO_DECISION, because the controller cannot prove what completed outside the transaction;screeningreturns topublishedorreproduction_pendingwith a retry disposition, preserving which lane was interrupted; and- a
leasedsettlement candidate returns topending, clears its lease, and increments the generation so a stale worker cannot commit it later.
A hold is not self-healing. Diagnose the retained reason, repair the authority, and use
the reviewed release/requeue API appropriate to the deployment. The store's
release_hold(...) appends an operator reason and chooses the lane from retained
publication and reproduction evidence; it does not erase prior attempts. No public CLI
wraps reservation-hold release, so deployment tooling must expose it under its
own access controls and audit trail.
Archive an exact schema-3 migration hold
One legacy database shape can retain a single-PASS schema-3 candidate that cannot satisfy the current two-PASS parser. It has a dedicated terminal operation:
cacheon chain-archive-schema3-hold \
--netuid <NETUID> \
--network <NETWORK_OR_WSS_URL> \
--intake-db chain_intake/intake.sqlite3 \
--reservation-id <RESERVATION_ID> \
--reason "reviewed migration reason"The command constructs no wallet. It accepts only the exact migration hold, records the current finalized height and bounded operator reason, preserves candidate and qualification bytes, removes the permanent queue veto, and can never release or crown the evidence. Generic expiry and hold release are not substitutes.
Incident playbook
| Alert | Immediate containment | Safe recovery criterion |
|---|---|---|
| Finalized cursor regression or changed hash | Stop the controller; preserve DB and endpoint logs | Chain endpoint/finality authority is understood; never overwrite the cursor |
| “another intake controller owns this database” | Find the legitimate owner; do not remove .lock | Exactly one live controller/signer window owns the DB |
| Repeated transport retry | Preserve URL, DNS, TLS, redirect, and archive evidence | Same committed bytes can be fetched within policy, or work remains held |
| Publication fault | Stop worker consumption of the affected address | Storage ownership/modes and independent reopen pass |
| Growing queue age | Stop new operational expansion; inspect screen/qualification capacity | Registered capacity and hardware can drain finalized order without reordering |
Cohort-wide NO_DECISION | Preserve failure digest and retry groups | Bisection or infrastructure repair completes under the same frozen authority |
| Evidence root unavailable | Block settlement and weights | Exact referenced artifacts reopen; rebuilding “equivalent” JSON is insufficient |
| Repeated pass exceptions | Let the bounded loop exit and quarantine the host if needed | Root cause fixed; one --once pass succeeds before daemon restart |
Operations checklist
- Put the database, private root, and publication root on durable local storage.
- Schedule
chain-snapshotagainst a private object-store namespace and runchain-snapshot-verifyafter every backup; copying only the main file while WAL writes are active is not a valid backup. - Test a retained fresh-root restore before relying on the archive, and keep bucket access policy, encryption, versioning/object lock, lifecycle, and capacity alerts under operator review.
- Alert on growing
held,no_decision, transport retry, and queue-age counts. - Monitor disk and inode use in both private and immutable publication roots.
- Run
cacheon chain-compatafter changing the Bittensor SDK. - Keep coldkeys off this host. Intake needs no wallet; the separate weight signer uses only the configured validator hotkey.
- Coordinate signer access between passes because the SQLite authority is single-owner.
- Retain service manifest, policy, logs, screen receipts, qualification artifacts, and software/image digests long enough to explain every standing claim.
Continue with Arena service and Settlement and weights.
Source anchors
Arena service
An arena service is the trusted bridge between an immutable proposal publication and crownable qualification. It binds what is being measured, where it may run, and how scarce evaluator capacity is allocated.
Qualification
Qualification asks a narrow question: does one exact submitted delta improve one frozen evaluation stack, in one registered arena, at acceptable quality?