Override points
Override points let a small miner-owned function fill a typed hole in a validator-owned base kernel. The resulting callable still implements a normal registered slot ABI; the base kernel, preparation path, output ownership, and serving integration remain validator-owned.
This is an advanced composition mechanism, not permission to patch arbitrary engine code.
Registered capability
The override registry defines one override-point identity:
moe.fused_experts/gemm1_epilogueIt selects the gemm1_epilogue hole in the validator-owned
nvfp4_moe_megakernel base. The registry and composition code are in
override.py.
Its support boundary is explicit:
- manifest parsing, point lookup, composition, and the dense/CPU reference path are implemented and tested;
- the override consumer raises
NotImplementedErrorwhen no validator-owned NVFP4 GPU provider exists; and - therefore the committed example supports ABI development but is not eligible for GPU qualification through this base.
Provenance measurements embedded in example metadata are not evidence produced by the override bundle.
Bundle declaration
A competitive version must still request its registered singleton target:
bundle_id = "my-moe-epilogue-v1"
abi_version = "cacheon-op-abi-v0"
[competition]
target = "moe.fused_experts"
mode = "slot"
[[ops]]
slot = "moe.fused_experts"
source = "kernels/epilogue.py"
entry = "gemm1_epilogue"
base_kernel = "nvfp4_moe_megakernel"
override_point = "gemm1_epilogue"
dtypes = ["bfloat16", "float16"]
metadata = "metadata/moe.json"The target catalog must allow the observed override feature, and the point
registry must resolve the (slot, override_point) pair. Manifest fields cannot
name a miner-supplied base kernel.
The paired callable convention
For entry = "gemm1_epilogue", the loader expects:
gemm1_epilogue_ref(gate, up, ...): a required Torch reference used by the dense path and fidelity checking;gemm1_epilogue(...): the optional-at-import device function, required for a real GPU implementation.
On a CPU host, guard CuTe/CUTLASS imports so the reference remains importable:
def gemm1_epilogue_ref(gate, up):
return activation(gate, up)
try:
import cutlass
import cutlass.cute as cute
@cute.jit
def gemm1_epilogue(t_compute, gate, up, alpha_val, *act_params):
...
except ImportError:
passIf the device implementation calls cute.compile(...), static closure accepts
that method only when cute is bound exactly once by import cutlass.cute as cute or the equivalent absolute from cutlass import cute. Rebinding or
shadowing the alias, importing it twice, using another module's .compile,
vendoring a local cutlass package, or passing a string/bytes literal as the
first positional argument fails closed. The validator-owned
DSL_JIT_ENTRYPOINTS table admits only CuTe. The exception exists only
for CuTe's function-object tracing API; it does not permit string-to-code
compilation or builtins.compile. Triton's lazy @triton.jit launch path does
not need an explicit .compile admission.
The authoritative admission logic is
dsl_jit_policy.py
and engine_tree.py.
The exact device ABI belongs to the registered point. For this point,
alpha_val is the per-expert quantization/dequantization value; activation
parameters such as a sigmoid gain are distinct values. Mixing those two
meanings is a correctness bug.
What composition does
The loader resolves the point, loads the required _ref function and optional
device function, and builds a standard fused_experts callable. In dense mode,
the validator-owned base implementation evaluates the miner's Torch activation.
In GPU mode, the base would install the device epilogue inside the megakernel.
The miner does not supply prepare for this override. The validator owns the
base kernel's weight layout and returns the composed (entry, prepare) pair.
Development checks
No public miner example advertises this point while its GPU provider is absent. The engine-tree test builds a minimal private fixture for composition coverage:
python -m pytest -q tests/test_cacheon_kernels.py tests/test_engine_tree.pyThis does not test the missing GPU base, graph replay, authoritative quality evidence, or performance.
When not to use an override
Use a normal slot implementation when you own the complete registered callable. Use a dependency patch only when a registered target explicitly requires an approved dependency export/build change. A hole, base, or engine change that is not registered cannot be submitted; registering it is a reviewed validator-side catalog change. Do not vendor a large dependency into a bundle to simulate a validator-owned base.
Troubleshooting
Diagnose a proposal from its last authoritative state. A local PASS cannot override a later intake, screen, qualification, or settlement result because each stage has different identity and evidence.
Dependency patches
A dependency patch is a narrowly allowed source delta against a pinned validator dependency. It exists for registered targets whose callable boundary depends on a small producer/export change in that dependency.