Cacheon

Bundle manifest

Every contribution bundle has a manifest.toml. The manifest is parsed as data before any contribution module is loaded. Its current ABI identifier is cacheon-op-abi-v0. Branding does not version this hash-bound protocol identifier, so Cacheon authors continue to emit the published spelling.

What parsing does—and does not do

A manifest passes through distinct layers:

load_manifest() covers only the first step: required syntax, identifier shape, path containment/existence, variant uniqueness, and structural CUDA or patch declarations. It does not import source, prove that the requested target exists, or verify numerical behavior. scan, target resolution, verify, and production qualification are separate gates so a structurally valid document cannot grant itself authority.

Minimal singleton example

bundle_id = "example-silu-v1"
abi_version = "cacheon-op-abi-v0"

[competition]
target = "activation.silu_and_mul"
mode = "slot"

[[ops]]
slot = "activation.silu_and_mul"
source = "kernel.py"
entry = "silu_and_mul"
dtypes = ["float32"]
architectures = ["cpu"]

bundle_id and identifiers accept letters, numbers, ., _, and -. Paths are relative to the bundle and must resolve to regular contained files.

Top-level fields

FieldRequiredMeaning
bundle_idyesHuman-readable bundle identifier; not the content identity
abi_versionyesNew bundles must emit cacheon-op-abi-v0; the reader also accepts the exact hash-bound pre-cutover spelling described below
[competition]recommendedExplicit requested target and slot/atomic mode
[[ops]]yesOne or more implementation rows
[[dep_patches]]noDeclared text patches for a validator-approved dependency lane

New bundles must use cacheon-op-abi-v0. Validators also retain one narrowly scoped reader-only compatibility spelling, optima-op-abi-v0, for finalized bundles committed before the rename. Readers preserve the submitted spelling and bytes; they do not rewrite a publication because the spelling participates in the committed bundle identity. Every other ABI spelling is refused.

The canonical bundle hash, not bundle_id, is the proposal's content identity.

Identity versus display name

The content hash walks the bundle's own regular files in sorted relative-path order and hashes length-prefixed path and byte sequences. Git and Python cache noise is excluded; symlinks are not part of the hashed file set and the bundle loader/scanner rejects unsafe tree structure. Editing source, metadata, a patch, or even the manifest produces a new identity. Renaming only the outer directory does not.

bundle_id is useful in logs and diagnostics but never proves that two reveals contain the same artifact. Commit/reveal, immutable publication, copy checks, and qualification use digest-bound content.

Competition table

FieldValues
targetA validator-registered target ID
modeslot or atomic; system parses only for legacy migration

The table is a request, not policy. Intake resolves it against the frozen target catalog and complete observed feature set. The current resolver can infer a target when an exact singleton or registered atomic member set is unambiguous, which preserves older bundles. New competitive bundles should declare [competition] explicitly.

Operation rows

FieldRequiredMeaning
slotyesRegistered execution slot
sourceyesPython source module within the bundle
entryyesEntry callable name
variantconditionalCapability variant; required on every row when a slot repeats
preparenoWeight-preparation callable for a registered prepare/forward ABI
setupdevelopment onlyLegacy parser/direct-framework diagnostic hook; every registered component target forbids it
dtypesnoDeclared dtype capability domain; an empty array adds no dtype restriction
architecturesnoDeclared architecture capability domain; an empty array adds no architecture restriction
metadatanoEligibility/capability JSON within the bundle
base_kernelnoValidator-owned base-kernel identifier for an override submission
override_pointnoTyped hole in that base; requires base_kernel
cuda_sourcesnoInspectable .cu/.cuh inputs for a reviewed builder
aot_exportsnoOrdered sealed direct-artifact exports for a slot with a declarative call ABI
artifact_resourcesnoValidator-allocated storage shared by the row's artifact lifecycle

Unknown operation fields are retained for observation. They do not grant a capability; complete feature resolution can reject a bundle that asks for more than its target admits.

How an operation row is selected

The slot chooses a validator-owned semantic ABI, not an arbitrary import hook. The validator resolves source inside the bundle and looks up the named Python identifier only after structural and static gates. It allocates outputs and passes arguments in the registered slot order. Candidate code fills those outputs; it does not redefine shapes, references, tolerances, or the call site.

For prepare/forward slots, prepare names the registered one-time weight transformation while entry names the runtime call. setup is a legacy direct framework hook and every current component target forbids it, so new competitive manifests should not use it.

For a row with aot_exports, source is scanned and its declared compiler factory is imported only in the isolated prebuild compiler child. The scheduler does not import that source as a runtime launcher. entry remains required by the outer cacheon-op-abi-v0 syntax, but direct execution canonicalizes it to the validator entry _cacheon_direct_artifact; changing the unused value does not change direct-execution identity.

Variants

Several rows may implement the same semantic slot only when every row has a unique explicit variant. Their validator-parsed capability domains must not overlap. Runtime selection is fail-closed: exactly one applicable variant is required, otherwise the trusted baseline is used or qualification refuses the candidate according to the registered boundary.

Empty or omitted dtypes and architectures do not mean “supports nothing”; they add no restriction at the manifest-parser layer. Eligibility metadata and validator observation still constrain the real capability domain. Conversely, listing sm103 does not prove the source builds or runs there.

[[ops]]
slot = "norm.rmsnorm"
variant = "sm90-bf16"
source = "rms_sm90.py"
entry = "rmsnorm"
dtypes = ["bfloat16"]
architectures = ["sm90"]

[[ops]]
slot = "norm.rmsnorm"
variant = "sm103-bf16"
source = "rms_sm103.py"
entry = "rmsnorm"
dtypes = ["bfloat16"]
architectures = ["sm103"]

This example demonstrates variant syntax only. norm.rmsnorm is unavailable for paid submission in the current MiniMax-M3 mainnet arena because that model does not execute the registered RMSNorm callsite. See Current MiniMax-M3 availability.

Sealed direct-artifact exports

[[ops.aot_exports]] is a strict, closed table. It is accepted only for a slot that has an immutable validator-defined artifact call ABI and only when target resolution permits the observed provider and rebuild features.

FieldRequiredContract
provideryesRegistered provider ID; the registry contains cutlass.cute.cubin.v1
nameyesExport-local canonical identifier
factoryyesIdentifier called only in the no-egress compiler child
profile_inputsyesUnique list drawn from the provider's compile-profile allowlist
bindingsyesOrdered projections of immutable slot resources and declared artifact resources
device_planyes for the registered providerComplete cacheon.device-launch-plan.v1 declaration
rolenoinit, prepare, reset, run, or destroy; default run
plannoSpecialization-plan name; default default
stepnoOrdered non-negative step within the plan; default 0
specializesnoExact-equality predicates over call-ABI resources
prelaunchnoBounded validator operation list; the registered operation is exact fill
provider_capability_requirementssealed/reopenCanonical requirements derived from binding projections
specialization_capability_requirementssealed/reopenCanonical requirements derived from specialization sources

The provider compile-profile allowlist is max_active_clusters.cluster_size_{1,2,4,8,16}. A factory can receive only the keys named in its profile_inputs; it cannot query a GPU in prebuild.

TOML row order is not executable authority. The parser canonicalizes exports by provider, plan, step, and name. Within each plan, steps must be unique and their step-sorted lifecycle roles must follow init, prepare, reset, run, destroy; every step in that plan carries the same specialization predicate. Plan overlap is legal only as a strict fallback chain. Runtime selects the unique most-specific exact match rather than using source order.

Binding rows

Each binding row has a kind and, except for aggregate, a canonical source such as input.q, output.out, stream.current, or a declared workspace.* resource. Available projections are a closed relation to the source kind:

Binding kindMain projections and options
tensordescriptor; optional bounded unsqueeze, checked assumed_align, and leading_dim
pointerdevice_ptr from tensor storage, identity from a pointer, or provider-authorized native_handle / peer_ptr_table from a group
scalarvalue, or tensor/group metadata shape, stride, numel, rank, size, storage_offset; requires an exact scalar cast
streamidentity for the validator's captured stream
groupidentity for the supplied semantic group
opaqueidentity for a validator-held opaque slot resource
aggregateBounded nested Shape, Coord, Tile, IntTuple, or Stride value with i32/i64 dynamic leaves

An axis is required for shape and stride. A peer_ptr_table binding must name the exact persistent group_ipc artifact resource in peer_resource. Bindings cannot name a callback, allocate a new stream or group, or construct an integer address.

Artifact resources

Each [[ops.artifact_resources]] row has exactly these authored fields:

FieldRequiredContract
nameyesworkspace.*, prepared.*, or state.*
dtypeyesRegistered storage dtype
alignmentyesPower of two, at least dtype width and at most 1 MiB
lifetimeyesFixed by the name prefix: call, prepared, or engine
shapeyesOne to 16 bounded extent tables
scopenorank_local by default or persistent group_ipc

Each shape extent contains factors and optional divisor. A factor is either { static = N } or a dynamic table with source, projection, upper_bound, and axis when projecting a tensor dimension. The extent is ceil(product(factors) / divisor). This is the entire allocation-expression language. Dynamic sources must be resources in the immutable slot call ABI, not other artifact buffers. A prepared-lifetime shape may use only resources present at the validator's exact prepare-allocation boundary.

Prefix, lifetime, and role are validated together:

PrefixLifetimeLifecycle rule
workspace.*callCannot be used by init or destroy
prepared.*preparedMust be produced in prepare and consumed in run; cannot be used by init or reset
state.*enginePersists with the artifact entry and is available to registered lifecycle roles

group_ipc is invalid for call-local workspace. Manifest ceilings are 32 resources, 64 GiB per resource, and 128 GiB aggregate; runtimes may impose lower live limits. These declarations request validator-owned storage, not candidate allocation authority.

Device launch plan

The registered provider requires device_plan.schema = "cacheon.device-launch-plan.v1" plus exact kernels and launches inventories.

InventoryRequired contents
kernelsSorted unique logical name and complete parameter_sizes byte vector
launchesContiguous ordinal, logical kernel, grid, block, optional cluster, shared_mem_bytes, ordered parameters, stream_binding, and allowlisted attributes

Grid, block, cluster, shared memory, and scalar parameter values are checked expression trees over admitted live bindings. Parameter kinds are pointer, scalar, packed_struct, tma_descriptor, cutlass_fast_divmod_i32_v1, cute_fast_divmod_i32_v1, and group_handle. The parameter-size vector of every launch must exactly match the logical kernel contract.

The plan schema is bounded to 1,024 logical kernels, 256 launches, and 64 semantic bindings. Dynamic shared memory is expression-derived; the runtime evaluates it for each invocation and rejects values above 1 MiB. At runtime, canonical logical kernel ordinal is matched to the complete driver-observed physical CUBIN inventory and exact per-ordinal widths. Candidate-provided physical-symbol search is not part of the ABI.

See Sealed direct artifacts for build, admission, evidence, identity, and support boundaries.

Dependency patches

[[dep_patches]]
target = "flashinfer"
path = "patches/change.patch"

Only UTF-8 .patch/.diff files are structurally accepted. Binary, rename, and deletion patches are refused. Admission still requires a target that permits the observed dependency-patch feature and a validator-reviewed applier with a bounded destination policy.

CUDA source declarations behave similarly. Listing .cu/.cuh paths makes them inspectable inputs to the sanctioned build lane; it is not permission to ship a prebuilt binary or execute an arbitrary compiler command. Rebuild operations come from reviewed validator policy.

What the manifest cannot do

A bundle cannot choose:

  • its reward family, overlap, or displacement;
  • arena hardware, workload, thresholds, or hidden tasks;
  • the incumbent, reference, or release stack;
  • isolation or network policy;
  • qualification, reproduction, or settlement outcomes.

Failure diagnosis

Error classTypical causeFix the right layer
TOML/required-field errorMissing [[ops]], wrong ABI string, invalid identifierCorrect manifest syntax
Path errorAbsolute path, .. escape, missing file, unsafe symlinkKeep every declared input as a regular contained file
Duplicate-slot errorMultiple rows without explicit unique variantsName every variant and make domains disjoint
Competition errorUnknown target, wrong slot/atomic mode, legacy system titleChoose a registered target from validator output
Feature-admission errorCUDA, patch, override, setup, or extra capability outside target policyRemove the feature or choose the registered lane
Artifact-provider errorUnknown provider, missing rebuild feature, or non-crownable providerUse the registered provider only on a target that admits it
Artifact-ABI errorBinding source/projection, resource lifecycle, or device parameter widths disagreeReconstruct the declaration from the slot call ABI and complete CUBIN contract
Compile-profile errorFactory requests an undeclared constant or architecture/profile authority differsDeclare only allowlisted inputs and rebuild for the exact arena profile
Static scan errorForbidden import/operation or uninspectable tree contentRewrite the source; scanning is not a sandbox exception list
Verification errorCallable/signature/output/reference mismatchDebug the registered tensor and correctness contract
Qualification failureComplete engine misses timing, drift, quality, fidelity, or resource gatesInspect retained arena evidence; do not relabel the outcome

Pre-submission checklist

  • Run slots and select the economic target separately from its execution slot members.
  • Declare [competition] explicitly for new singleton or atomic work.
  • Keep bundle_id descriptive but assume only the content hash is identity.
  • Declare every source, metadata, CUDA, and dependency-patch input with a contained relative path.
  • Give repeated slot rows unique variants with non-overlapping domains.
  • Avoid setup; use only registered prepare and entry contracts.
  • For direct artifacts, declare the complete binding, resource, lifecycle, and device plan; never package generated CUBIN or a host launcher.
  • Run scan, then the appropriate local or collective verify command.
  • Hash and package the exact verified tree; any later byte change is a new proposal.

Source: cacheon/manifest.py.

On this page