Cisco Caceres · self-published independent research note · version 1 · 7 October 2026 Pacific
Status and scope
This note proves a small, standard-ingredient statement about choosing one answer from a fixed pool of candidates when the only information is a set of verifier signals. It is a mathematical note. It contains no model results, no frontier-model claim and no claim of novelty: every inequality used is textbook (Jensen, Pinsker, the chain rule for mutual information, the optimality of Bayes decisions). It is not peer reviewed. The statements, proofs and finite checks were examined by a separate AI agent working in the same research workflow; that is a model-assisted check, not outside review and not independent reproduction.
The proposed 24-task verifier study described on the research proposal page has not been run, and this note changes none of its protocol. What is proved here applies to the fixed-pool setting defined below and nowhere else; the last section lists the ways it fails to transfer to real selection policies.
Setting and assumptions
- A1, fixed pool. One task, K ≥ 2 candidates fixed before any verifier is consulted. Candidate labels Y = (Y₁,…,Y_K), each 0 or 1, are unknown when the choice is made.
- A2, exogenous signals. S is the full vector of verifier outputs that could be collected for this pool. Its values do not depend on the order in which a policy asks for them. The pair (Y, S) has a joint law P, treated as known.
- A3, policies. A policy chooses one index as a function of the coordinates of S it has observed and of independent internal randomness. It may choose adaptively which coordinates to observe.
Notation: pᵢ = P(Yᵢ = 1); qᵢ(s) = P(Yᵢ = 1 | S = s); C = P(some Yᵢ = 1), the candidate coverage (pass@K); B₀ = maxᵢ pᵢ, the best choice with no information; I(·;·) is mutual information in nats. The information-limited value is V*(S) = E[ maxᵢ qᵢ(S) ].
Proposition 1: the information-limited ceiling
- Every policy satisfying A3 has expected selected correctness at most V*(S), and V*(S) ≤ C.
- V*(S) ≥ B₀.
- If S′ is a function of S, then V*(S′) ≤ V*(S).
- V*(S) = C exactly when, for every signal value s with positive probability, maxᵢ P(Yᵢ=1, s) = P(some Yᵢ=1, s).
- If A and B are two blocks of coordinates of S and every Yᵢ is independent of B given A, then V*(A, B) = V*(A).
Proof. (1) A policy whose decision depends on S and on randomness U independent of (Y, S) has expected correctness Σᵢ E[ P(choose i | S) qᵢ(S) ], which is at most E[maxᵢ qᵢ(S)] because the weights form a probability vector for each s. Adaptive acquisition only makes the decision a function of the coordinates it acquired, hence of S, because A2 says the values do not depend on the order of acquisition. Next, V*(S) = Σₛ maxᵢ P(Yᵢ=1, s) ≤ Σₛ P(some Yᵢ=1, s) = C termwise. (2) By Jensen for the convex maximum, E[maxᵢ qᵢ] ≥ maxᵢ E[qᵢ] = B₀. (3) For S′ = f(S): V*(S′) = Σₛ′ maxᵢ Σ over s in f⁻¹(s′) of P(Yᵢ=1, s) ≤ Σₛ′ Σ over s in f⁻¹(s′) of maxᵢ P(Yᵢ=1, s) = V*(S). (4) The termwise inequality in (1) is an equality for all s exactly when V* = C. (5) Under the independence assumption qᵢ(a, b) = qᵢ(a), so V*(A, B) = E[maxᵢ qᵢ(A)] = V*(A). ∎
Reading. For any baseline policy π₀, C − E[correctness of π₀] = (C − V*) + (V* − E[correctness of π₀]). The first term is information deficit: what the signals cannot reveal even in principle. Only the second term is recoverable by a policy that uses these signals. The coverage oracle C alone therefore overstates what any verifier-based selector can attain. V* is a ceiling for forced single-choice accuracy only.
Proposition 2: an information bound on the attainable gain
Let G = V*(S) − B₀, the most any policy can gain over the best no-information choice. Then
G ≤ √( ½ · Σᵢ I(Yᵢ ; S) ).
By the chain rule, I(Yᵢ ; S) splits across any partition of the signals, for example I(Yᵢ ; S) = I(Yᵢ ; A) + I(Yᵢ ; B | A) for two verifiers A and B. The bound is therefore a sum of a first-verifier term and a conditional second-verifier term. If the conditional term is zero for every candidate, Proposition 1(5) shows the second verifier adds nothing to V*.
Proof. Fix s. Since Yᵢ is binary, qᵢ(s) − pᵢ is at most the total variation distance between P(Yᵢ | S=s) and P(Yᵢ). Pinsker's inequality, in nats, bounds that distance by √( KLᵢ(s) / 2 ), where KLᵢ(s) is the Kullback–Leibler divergence of P(Yᵢ | S=s) from P(Yᵢ). Hence maxᵢ qᵢ(s) − maxᵢ pᵢ ≤ maxᵢ √(KLᵢ(s)/2) ≤ √( ½ Σᵢ KLᵢ(s) ). Take the expectation over S and use Jensen for the concave square root: G ≤ √( ½ Σᵢ E[KLᵢ(S)] ), and E over S of KLᵢ(S) equals I(Yᵢ ; S). ∎
Remarks. The bound is not tight. On every exhaustively enumerated model below, the actual gain is at most 0.70 of the bound. It is vacuous when ½ Σᵢ I ≥ 1. It uses the whole pool's signals for each label, so it includes what other candidates' signals say about candidate i. It is a ceiling on gain, not an estimate of it.
Proposition 3: pairwise correlation and effective vote counts do not determine the ceiling
Statement. There are two joint laws of one candidate's label and a pair of ternary verifier outputs (A, B) with identical conditional marginals and identical conditional covariance under each label, hence identical correlation and identical effective-vote-count summaries, for which a two-candidate pool has different V* and different total information.
Exhibit. Two independent candidates each have prior P(Y = 1) = 3/10. Each candidate's pair (A, B) takes values 0, 1, 2 with these cell counts out of 24 (rows are A, columns are B):
| Law | Label | Cell counts |
|---|---|---|
| 1 | Y = 0 | (10, 2, 2), (2, 4, 0), (2, 0, 2) |
| 1 | Y = 1 | (0, 2, 2), (2, 4, 0), (2, 0, 12) |
| 2 | Y = 0 | (11, 0, 3), (3, 2, 1), (0, 4, 0) |
| 2 | Y = 1 | (0, 4, 0), (1, 2, 3), (3, 0, 11) |
Under Y = 0 both laws have row and column sums (14, 6, 4) and Σ a·b·count = 12. Under Y = 1 both have row and column sums (4, 6, 14) and Σ a·b·count = 52. These are all the moments that fix marginals and covariance. Yet exact arithmetic gives V* = 69/160 = 0.43125 for law 1 and V* = 2413/4800 ≈ 0.50271 for law 2, against coverage C = 51/100 and a no-information baseline of 3/10. Total information Σᵢ I(Yᵢ ; S) is 0.3638 nats for law 1 and 0.9532 nats for law 2.
Why. Selection uses the ordering of posterior probabilities over the nine signal cells, and that ordering depends on the whole table rather than on its first two moments. For two binary verifiers the conditional table is fixed by the marginals and the correlation, so a counterexample of this kind needs at least three signal levels, or dependence across candidates.
A known-law contrast with additive scoring
Call a rule additive if it picks the candidate maximizing g_A(aᵢ) + g_B(bᵢ). For one pair of binary verifiers an exact search over all integer-weight conditional laws with 10 units of weight found the largest gap between V* and the best additive rule: with Y = 0 cells (A,B) = (0,1) and (1,0) each weighing 5 and Y = 1 cells (0,0) and (1,1) each weighing 5, V* = 51/100 = C while no additive rule exceeds the no-information value 3/10. This is an adversarial XOR-like law shown to prove that the gap can be as large as the whole attainable headroom, not to suggest it is typical. For law 1 above an additive rule reaches V* exactly (gap 0); for law 2 the best additive rule found on an integer grid reaches 0.445 against V* = 0.503, an upper bound on the gap only.
This is a known-law, fixed-pool contrast between the Bayes-optimal rule and the best additive scoring rule, with no budget. It is not the planned budgeted comparison of a joint-error policy with an additive policy, and no budgeted comparison has been run or bounded here.
Where the ceiling does not transfer to a real policy
V* bounds a real policy only if every one of the following holds. Each can fail, and none is claimed.
- Information set. Every signal the policy uses must be a coordinate of S. A policy that obtains information outside S is not bounded. Example: Y₁, Y₂ independent fair coins and S uninformative gives V* = 1/2, while a policy that also sees both labels reaches 3/4.
- Exogeneity. Signals must not depend on the policy's actions. Interactive or stateful verifiers, repeated sampling with feedback, or tool use that changes the output violate A2.
- Fixed pool. Regenerating or extending candidates changes the pool and C. The ceiling does not bound best-of-n with larger n.
- Objective. V* concerns forced single-choice accuracy. Abstention or selective accuracy is a different objective and is not a function of V*: two laws with the same V* = 3/4 can have coverage at precision 0.9 of 1/2 and 0.
- Cost and budget. V* ignores query cost. It is an upper limit for any budgeted adaptive policy that uses these signals, but the note gives no tighter budgeted bound.
- Law and estimation. The statements hold for the true law P. A learned policy trained on finite data is not bounded by a data-independent expression here, and the ceiling cannot be computed without P.
- Scope. Binary labels, one task, one fixed pool. Averaging over tasks preserves the inequalities by linearity of expectation only if the bound is applied per task.
Finite checks and reproduction
A standalone script enumerates every integer-weight joint table of two labels and a signal on small grids, with exact arithmetic for V* and C and a 1e-12 float slack for the information bound. It found no violation of Proposition 1 or 2 over 141,474 tables (signal alphabets of 2, 3 and 4 symbols), confirmed the garbling and conditional-independence statements, and reproduced the exhibit above. These checks are evidence, not proof: the models are small, K = 2 and on a grid. The proofs above are the argument.
The package contains the checking script, its output, the small exact examples and fast regression tests, and needs only Python 3's standard library. The complete enumeration takes a few minutes; the regression tests take a fraction of a second.
What this note does not establish
No closed form, lower bound or general bound for the contrast between a joint and an additive policy at a budget is derived. Nothing here is about any real verifier, code model or benchmark. The 24-task study is unrun. The related-work context (correlated verifier errors, imperfect verifiers) is in the proposal; this note asserts no relationship to it beyond the setting.
Downloads
The note (Markdown) · Checking script, outputs and tests (ZIP; Python standard library only) · Source hashes.