Deterministic LoRA training directly through quantized GGUF serving weights

Thank you. That answers everything, and more carefully than i asked.

The licence settles it: the 36 prompts and the family grouping will come into Runner under Apache-2.0 with your name on them, as the seed corpus for the native boundary lane, and the bank repo will be referenced as the canonical source. The clean and executed notebooks stay linked as the discovery record. You are right that the native lane should not erase how the result was found, and the discarded sequence scorer is exactly the kind of history worth keeping: it documents why the lane records grammar branches instead of tool-name text.

Your provenance framing is adopted verbatim, and it is now written into the planning record in your words so that a future reproduction report cannot quietly make the result stronger than what was run: published training provenance, to exact published adapter bytes, to independent behavioral probe, to confirmed full-call decision flip. That chain is closed on the published study objects. It is not an independent three-precision training reproduction, and nothing we publish will describe it as one. The T4 one-step byte-identity result stays filed as what it is, a separate experiment about the artifact-determinism contract.

The migration bridge is adopted as you designed it. First run of the native lane will be your seven cases against exactly the artifacts you pinned, the three sha256 adapters over the shared Q4 serving base, through choice_logprobs. The invariant we will hold it to is the behavioral one: same artifacts, same prompts, same qualitative branch separation, same full-call disagreement. We will not require the native margins to reproduce the notebook’s first-token margins, because the two constructions measure different surfaces, and if a boundary case moves or disappears under the native grammar we will record that as protocol dependence rather than as either measurement being wrong.

If the bridge passes, the next step is the one you suggested: all 36 prompts under base, BF16, Q8 and Q4 with no adapter-screened subset, reported as a property of this bank and nothing more. Your three-quantity separation is now standing wording on our side: one of seven selected cases is an existence proof, disagreement across the 36 constructed prompts is a property of the bank, and prevalence in real traffic would need a sampling design nobody has built yet. The exploratory lane stays unlabeled, with the record shape you listed. If a calibration lane is ever worth building it will be a separate dataset with labels frozen before any adapter output is seen, feeding the existing calibration script, and it will not replace the exploratory bank.

When the boundary lane lands in Runner it will carry the attribution and a pointer back here, and the changelog will say where the prompts came from.

For whatever it is worth from the other side of the exchange: preregistered screens, instruments you killed yourself, and claims scoped tighter than i would have dared to scope them for you.

Thank you again.