I am developing an independent AI research program alongside my work in AI systems architecture. I want to understand how models reason, how we judge their answers, and how to spend compute where it improves reliable task completion.
My direction includes research relevant to frontier models: inference-time reasoning and evaluation first, followed by controlled model adaptation and learning experiments. Each contribution will need a clear question, a comparison against existing work, and reproducible evidence.
Correlated verifier errors and calibration drift in budgeted code generation. If several judges share the same mistake, agreement can create false confidence. I plan to compare selection and stopping decisions against execution-based controls and published methods, with every inference action charged to the budget.
Each project has its own evidence and next experiment. Research includes proposals, design work and empirical investigation; their stages remain distinct.
Research proposal · Engineering feasibility measuredQuestion: How do correlated verifier errors and calibration drift affect budgeted code-generation decisions?
Method: Licensed task freeze and isolated executable development checks; planned matched-controller comparison with hidden tests, fixed budgets and separate natural-shift and synthetic-error arms.
Artifacts: Updated proposal, related-work matrix, frozen HumanEval splits, guarded local collector and networkless VM development qualification.
Evidence: V2 retains fourteen response-contract failures; V3 has twelve public/hidden passes on six development tasks. This saturated set establishes no verifier-selection benefit, specificity, study hypothesis result or novelty.
Next milestone: Run the separately frozen 120-task pilot with matched-token baselines and retain all contract failures and costs.
Read the evidence and method
Model research · Live commercial-engine serviceQuestion: Which acoustic and deployment choices make noisy-speech recognition more reliable under radio and environmental constraints?
Method: Existing third-party ASR comparison under clean and synthetic babble conditions; planned licensed, speaker-separated channel and noise evaluations for proprietary models.
Artifacts: Dated third-party speech benchmark, tested architectural assumptions and a live commercial-engine passthrough service. A frozen speaker-separated local third-party ASR diagnostic and insertion-heavy failure analysis.
Evidence: The published benchmark evaluates third-party models. It establishes no AMBIE-trained model accuracy. Four of five initial architectural assumptions were falsified.
Next milestone: Extend the frozen acoustic protocol to licensed domain-relevant channels and critical fields before assessing a proprietary checkpoint.
Read the evidence and method
Whitepaper and design · Implementation pendingQuestion: What does machine intelligence cost in energy, and which proposed computational mechanisms can be tested against physical constraints?
Method: Proposed continuous-time dynamics studies, followed by a digital implementation and comparisons with conventional computation before any hardware experiment.
Artifacts: Written whitepaper and design materials; public project description.
Evidence: No software or hardware implementation and no measured results. Proposed energy figures are not demonstrated performance.
Next milestone: Define a minimal falsifiable experiment, implement it and measure it against a conventional baseline.
Experimental acoustic research · Internal evaluationQuestion: Can interpretable acoustic features help distinguish synthetic speech across generators, speakers and recording conditions?
Method: Complete-recording DSP analysis with selected internal stress tests; independent, speaker- and generator-separated evaluation is next.
Artifacts: Browser demo, feature analysis, internal test methodology and documented failure cases. A frozen human-speech babble diagnostic with retained false alarms.
Evidence: Scores are uncalibrated. Internal convenience samples include substantial In-the-Wild failures; verdicts do not authenticate speakers. Streaming integrations remain proposed designs. The human-only diagnostic estimates no synthetic miss rate or binary accuracy.
Next milestone: Audit labels and licensing, freeze splits and thresholds, then compare against learned and simple acoustic baselines.
Read the evidence and method
Operating research infrastructure · Documented ML experimentsQuestion: How can specialist-model training, adaptation and evaluation improve task outcomes within measured compute budgets?
Method: Specialist task-method comparisons, adaptation and compression ablations, speech and diagnostic benchmarks, and bounded single-node multi-GPU trials with per-arm evaluation and cost accounting.
Artifacts: Training-run tooling, dated provisioning ledger, model experiment matrices, ablation reports and per-arm metrics, including the former VastOps work. Exact saved-output rescore, recovered ablation provenance, local verifier diagnostics and a checked mathematical reference.
Evidence: Completed experiments retain task, dataset, simulator and hardware limits. The reported structural-pruning/recovery arm was planner-refused before training. Single-node multi-GPU execution does not establish multi-node or arbitrary-scale training; provisioning samples are not marketplace-wide rates.
Next milestone: Validate specialist findings on independent real data and held-out domain shifts with a frozen protocol.
Read the evidence and method
These working systems enable experiments or explain mechanisms. Their engineering evidence stays separate from the findings of studies run with them.
Pre-release coordinationSupports future studies of coordination, recovery and task completion in AI coding teams. Agent messaging, leased tasks, persisted results and advisory file reservations.
Artifacts: CLI/daemon implementation, transport measurements and recovery qualification records. An isolated source fixture with nine actual CLI checks and eight selected source qualifications.
Evidence: Transport performance and result-record integrity do not establish task correctness or comparative task-success gains.
Next milestone: Execute the prepared equal-resource coding comparison with pinned tasks, acceptance criteria and fault injection.
Working physical-machine toolingSupports studies of observable physical-machine setup, recovery and agent interventions. Out-of-band KVM screen reading and HID input alongside exact text and network checks where available.
Artifacts: Documented ARM-board bring-up, older-PC installation and x86 rescue sessions; Libre Computer setup also recorded.
Evidence: Each ARM and x86 record states its method and limits. The Libre card was prepared on another host, not through Bootscry delivery. Hangs and manual power interventions are retained; these are not a general recovery success rate.
Next milestone: Define repeated hardware trials with intervention counts, verified end states and a scripted baseline.
Educational implementationSupports understanding of network learning, tokenization, backpropagation and inference mechanics. Interactive small-model implementations in the browser.
Artifacts: Working educational demonstrations. A separate causal-attention educational reference passed nine tests on Haywire and Feral; numerical gradient and KV-cache checks are recorded.
Evidence: Teaching implementations do not establish a novel model method or frontier-scale training contribution.
Next milestone: Use a separately versioned training experiment for any future learning or scaling claim.
Versioned protocols, exact model and dataset identities, code, cost ledgers, uncertainty intervals, ablations and failure cases. I will distinguish conceptual replications from original contributions, publish negative findings, and seek independent reproduction and human review.
The research agenda and study status live here. Long-form technical reporting also belongs in the RoamingPigs Field Manual.
I welcome specific criticism of the protocol and collaborators interested in verification, code generation and reproducible inference research. Contact me about the study.