Methodology & claims
What the current system measures—and what it does not.
The current technical alpha describes raw chosen-minus-rejected differences on resolved, version-pinned interpretability primitives. It is explicitly unrated and does not make a “likely to move,” confidence, or training-outcome claim. The separate public claim path remains unavailable until calibration and evidence gates pass.
Measured
Trusted operator-side workers evaluate declarative dataset pairs with pinned model, runtime, hook, artifact, catalog, aggregation-policy, and report-schema bindings.
Disclosed
Reports preserve excluded rows, unreadable concepts, runtime failures, per-family reads, contested evidence, provenance, advisories, and the frozen manifest.
Not claimed
A diagnostic is not a guarantee of model behavior, training outcome, safety, benchmark performance, or business impact.
Explicitly unresolved
- The flat price is versioned and admin-managed; dataset byte/row caps await the compute benchmark.
- Launch catalog, artifact licensing, and catalog gate thresholds await approval.
- Calibration-tier meanings and a confidence estimator are not defined.
- Within-family independence and cross-family corroboration policy are not assumed.
- “Subnet-validated” positioning is not used by the current product.
The public sample is synthetic and unrated. It exists to prove contract shape and rendering, not to supply calibration, economics, novelty logic, or a real diagnostic.
Inspect the synthetic sample