Cisco Caceres · Proposal v0.2 · 6 October 2026
Status: protocol draft. Local development diagnostics have run; matched-controller study results, novelty and publication status are unestablished. Preregistration is planned and has not occurred.
When code-generation systems spend more compute choosing an answer, what happens if their verifiers share mistakes or their confidence stops matching correctness?
Execution evidence
The first response contract retained fourteen formatting failures. A separately frozen structured-module follow-up produced twelve candidates on six development tasks that passed public and hidden execution checks in an owned networkless VM. This all-positive set establishes neither verifier discrimination nor a selection benefit. The separately frozen 120-task engineering pilot uses token budgets; it is not the proposed natural-shift confirmation study or its dollar endpoint. Read the dated methods, failures and limitations.
The question
I plan to measure correlated verifier errors and calibration drift under a fixed inference budget. The candidate contribution is an empirical characterization and reproducible decision benchmark. Adaptive compute and imperfect-verifier limits are established research topics; this proposal does not claim to invent them.
A falsifiable hypothesis
On a held-out task-family shift, a controller calibrated to joint verifier errors will improve hidden-test pass rate over the same controller with joint-error interaction features removed, at one fixed budget. The practical design target is a three-percentage-point improvement, not a predicted result. A paired confidence interval will determine whether the evidence supports improvement, is inconclusive, or contradicts it.
Agreement is not proof: judges can agree on the same wrong answer. A disagreement-only stopping rule will be an ablation. Synthetic correlated-error interventions will probe mechanisms separately from the primary natural-shift evidence.
Experiment design
- One pinned open-weight generator, two verifier families, public-test execution, and up to sixteen cached candidates per task. Live feedback-conditioned revisions are outside the first replay experiment.
- Source-family-only development and calibration, with separate pilot and confirmatory target-family tasks; deduplication by source and problem family. The proposed pilot is 120 tasks. Confirmatory size follows a power calculation, rather than an arbitrary sample count.
- Controllers see public tests and requested verifier judgments. Hidden outcomes are used only for offline evaluation; future cached scores are unavailable to the controller.
- Audit augmented hidden tests and report correctness against the declared suite, candidate coverage, false acceptance, discrimination and calibration. Passing tests does not prove general correctness.
- Compare single-pass and first-public-test-pass controls, extra sampling, fixed-budget best-of-N, execution agreement, calibration-only stopping, joint-error-aware selection and an applicable published robust-selection method.
- Freeze policies, thresholds, task manifests, model revisions, primary comparison and analysis before confirmatory outcomes are inspected.
Compute and analysis
The primary cap will be four times median single-pass inference cost on the development split, using a frozen price sheet. Two secondary caps explore sensitivity. Charge generation, verification, repeated context, retries and failures; also report tokens, CPU or GPU time and latency. Billed cost is an operating metric, not a FLOPs measurement.
Replay charges only actions a policy requests. The full cached pool remains a separate actual research expense. Natural shifts can change task difficulty and candidate coverage as well as calibration, so they will not be presented as a causal isolation of drift.
The primary comparison keeps the estimator, other features and stopping rule matched. Freeze outcome suites before viewing controller results. Inference is conditional on one candidate pool and one family shift; pool exhaustion and later audit corrections will be reported. The primary endpoint is paired task-level hidden-test success, with a 95% confidence interval. Abstention counts as failure there and is reported separately in risk/coverage analysis. Secondary tests will account for multiple comparisons. Null findings and inconclusive intervals will be published.
Prior work and overlap
The literature already includes adaptive allocation, reward hacking, imperfect coding verifiers, distribution-shift-aware judges and execution-based selection. A full-text and code review must establish which comparisons are applicable before this study proceeds.
- Snell et al.: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. Adaptive inference allocation; conceptual replication only, not reproduction of inaccessible original models.
- Stroebl, Kapoor and Narayanan: Inference Scaling fLaws: The Limits of LLM Resampling with Imperfect Verifiers. Direct overlap: imperfect coding tests limit resampling. This study must go beyond that established finding.
- Huang et al.: Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment. Pessimistic selection under imperfect rewards; candidate strong baseline, assumptions need matching.
- Khalaf et al.: Inference-Time Reward Hacking in Large Language Models. Best-of-Poisson and HedgeTune; mandatory feasibility assessment for robust-selection baselines.
- Montgomery et al.: Budget-aware Test-time Scaling via Discriminative Verification. Compute-efficient verifier/self-consistency hybrid; not a new research direction here.
- Li et al.: S*: Test Time Scaling for Code Generation. Execution-grounded pairwise selection and distinguishing inputs; code-specific comparison.
- Qin et al.: DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation. Direct shift overlap: difficulty, task distribution and generator-trajectory mismatch.
- Qu: Adaptive Test-Time Compute Allocation via Learned Heuristics over Categorical Structure. Selective intermediate verification; structured-math assumptions may not transfer to final code selection.
- He et al.: Code Generation by Differential Test Time Scaling. Execution clustering without additional selection-model calls; include CPU costs.
- Sriraman and Block: Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment. Objective matters: win-rate and correctness are different endpoints; assess tuned BoN variant.
Full-text overlap review
ADAP already learns stopping and draws on ground-truth calibration; public-test proxies cannot inherit its guarantee. CAPS already budgets pairwise judging and eliminates finalists; scalar-score selection does not reproduce it. DAJ addresses judge distribution mismatch with learned data reweighting. GRACE studies verification granularity under assumptions not automatically satisfied by correlated code-judge errors. The narrower proposed contrast is conditional cross-verifier error dependence after matching marginal calibration, candidate coverage, action information and total cost. That difference remains a question to test, not a proven original contribution.
Schedule and release
The licensed HumanEval manifest, initial replay and isolated development checks are prepared. Next: run the separately frozen engineering pilot, finish applicable strong-method baselines, define genuine task-family shift, price the powered study, and archive an immutable preregistration. Confirmatory execution follows only if the question and budget survive those gates.
Planned artifacts include protocol versions, licensed task manifests, controller code, model settings, seeds, sanitized traces, cost records, analysis and failure cases. Model-assisted critiques informed this draft; they are not human peer review. No model-training run or new foundation-model architecture is claimed.
A separate transfer study would be needed to establish relevance across stronger models and other benchmarks. This proposal is a first step toward a broader model-research program.