AI Research Agenda

AI research by Cisco Caceres: AMBIE, Noise or Voice, ThermoCog, TuneHarness and verifier reliability. Questions, methods, evidence and next experiments.

I am developing an independent AI research program alongside my work in AI systems architecture. I want to understand how models reason, how we judge their answers, and how to spend compute where it improves reliable task completion.

My direction includes research relevant to frontier models: inference-time reasoning and evaluation first, followed by controlled model adaptation and learning experiments. Each contribution will need a clear question, a comparison against existing work, and reproducible evidence.

Portfolio status · 6 October 2026 (UTC): this agenda connects acoustic research, synthetic-speech detection, a design-stage energy project, documented specialist-model experiments and a new verifier-reliability proposal. TuneHarness includes experiments run under its former VastOps name. The new verifier study has not run; completed experiments retain their own methods and limits.

New proposal: when the verifier is wrong

Correlated verifier errors and calibration drift in budgeted code generation. If several judges share the same mistake, agreement can create false confidence. I plan to compare selection and stopping decisions against execution-based controls and published methods, with every inference action charged to the budget.

Research projects

Each project has its own evidence and next experiment. Research includes proposals, design work and empirical investigation; their stages remain distinct.

Research proposal · Engineering feasibility measured

Verifier reliability

Question: How do correlated verifier errors and calibration drift affect budgeted code-generation decisions?

Method: Licensed task freeze and isolated executable development checks; planned matched-controller comparison with hidden tests, fixed budgets and separate natural-shift and synthetic-error arms.

Artifacts: Updated proposal, related-work matrix, frozen HumanEval splits, guarded local collector and networkless VM development qualification.

Evidence: V2 retains fourteen response-contract failures; V3 has twelve public/hidden passes on six development tasks. This saturated set establishes no verifier-selection benefit, specificity, study hypothesis result or novelty.

Next milestone: Run the separately frozen 120-task pilot with matched-token baselines and retain all contract failures and costs.

Read the evidence and method

Model research · Live commercial-engine service

AMBIE

Question: Which acoustic and deployment choices make noisy-speech recognition more reliable under radio and environmental constraints?

Method: Existing third-party ASR comparison under clean and synthetic babble conditions; planned licensed, speaker-separated channel and noise evaluations for proprietary models.

Artifacts: Dated third-party speech benchmark, tested architectural assumptions and a live commercial-engine passthrough service. A frozen speaker-separated local third-party ASR diagnostic and insertion-heavy failure analysis.

Evidence: The published benchmark evaluates third-party models. It establishes no AMBIE-trained model accuracy. Four of five initial architectural assumptions were falsified.

Next milestone: Extend the frozen acoustic protocol to licensed domain-relevant channels and critical fields before assessing a proprietary checkpoint.

Read the evidence and method

Whitepaper and design · Implementation pending

ThermoCog

Question: What does machine intelligence cost in energy, and which proposed computational mechanisms can be tested against physical constraints?

Method: Proposed continuous-time dynamics studies, followed by a digital implementation and comparisons with conventional computation before any hardware experiment.

Artifacts: Written whitepaper and design materials; public project description.

Evidence: No software or hardware implementation and no measured results. Proposed energy figures are not demonstrated performance.

Next milestone: Define a minimal falsifiable experiment, implement it and measure it against a conventional baseline.

Experimental acoustic research · Internal evaluation

Noise or Voice

Question: Can interpretable acoustic features help distinguish synthetic speech across generators, speakers and recording conditions?

Method: Complete-recording DSP analysis with selected internal stress tests; independent, speaker- and generator-separated evaluation is next.

Artifacts: Browser demo, feature analysis, internal test methodology and documented failure cases. A frozen human-speech babble diagnostic with retained false alarms.

Evidence: Scores are uncalibrated. Internal convenience samples include substantial In-the-Wild failures; verdicts do not authenticate speakers. Streaming integrations remain proposed designs. The human-only diagnostic estimates no synthetic miss rate or binary accuracy.

Next milestone: Audit labels and licensing, freeze splits and thresholds, then compare against learned and simple acoustic baselines.

Read the evidence and method

Operating research infrastructure · Documented ML experiments

TuneHarness

Question: How can specialist-model training, adaptation and evaluation improve task outcomes within measured compute budgets?

Method: Specialist task-method comparisons, adaptation and compression ablations, speech and diagnostic benchmarks, and bounded single-node multi-GPU trials with per-arm evaluation and cost accounting.

Artifacts: Training-run tooling, dated provisioning ledger, model experiment matrices, ablation reports and per-arm metrics, including the former VastOps work. Exact saved-output rescore, recovered ablation provenance, local verifier diagnostics and a checked mathematical reference.

Evidence: Completed experiments retain task, dataset, simulator and hardware limits. The reported structural-pruning/recovery arm was planner-refused before training. Single-node multi-GPU execution does not establish multi-node or arbitrary-scale training; provisioning samples are not marketplace-wide rates.

Next milestone: Validate specialist findings on independent real data and held-out domain shifts with a frozen protocol.

Read the evidence and method

Planned model-method studies

  • Inference-time reasoning: how generation, verification and stopping interact under fixed budgets. First proposal published; experiments pending.
  • Evaluation and reliability: hidden-test quality, correlated judgment errors, distribution shift and uncertainty. These are questions to measure, not solved problems.
  • Efficient model adaptation: planned controlled fine-tuning and ablations against an unchanged base model. No results claimed.
  • Learning and scaling: planned small-model training studies that expose data, optimization and compute tradeoffs. Small-model evidence will stay separate from frontier-scale claims.
  • Speech and acoustic intelligence: an existing applied specialty and an ongoing research direction; the dated speech benchmark records what changed my assumptions.

Research infrastructure and education

These working systems enable experiments or explain mechanisms. Their engineering evidence stays separate from the findings of studies run with them.

Pre-release coordination

Echo

Supports future studies of coordination, recovery and task completion in AI coding teams. Agent messaging, leased tasks, persisted results and advisory file reservations.

Artifacts: CLI/daemon implementation, transport measurements and recovery qualification records. An isolated source fixture with nine actual CLI checks and eight selected source qualifications.

Evidence: Transport performance and result-record integrity do not establish task correctness or comparative task-success gains.

Next milestone: Execute the prepared equal-resource coding comparison with pinned tasks, acceptance criteria and fault injection.

Working physical-machine tooling

Bootscry

Supports studies of observable physical-machine setup, recovery and agent interventions. Out-of-band KVM screen reading and HID input alongside exact text and network checks where available.

Artifacts: Documented ARM-board bring-up, older-PC installation and x86 rescue sessions; Libre Computer setup also recorded.

Evidence: Each ARM and x86 record states its method and limits. The Libre card was prepared on another host, not through Bootscry delivery. Hangs and manual power interventions are retained; these are not a general recovery success rate.

Next milestone: Define repeated hardware trials with intervention counts, verified end states and a scripted baseline.

Educational implementation

LLM Lab

Supports understanding of network learning, tokenization, backpropagation and inference mechanics. Interactive small-model implementations in the browser.

Artifacts: Working educational demonstrations. A separate causal-attention educational reference passed nine tests on Haywire and Feral; numerical gradient and KV-cache checks are recorded.

Evidence: Teaching implementations do not establish a novel model method or frontier-scale training contribution.

Next milestone: Use a separately versioned training experiment for any future learning or scaling claim.

Read or download the research and systems portfolio for the contribution record and planned publication sequence.

What will count as evidence

Versioned protocols, exact model and dataset identities, code, cost ledgers, uncertainty intervals, ablations and failure cases. I will distinguish conceptual replications from original contributions, publish negative findings, and seek independent reproduction and human review.

The research agenda and study status live here. Long-form technical reporting also belongs in the RoamingPigs Field Manual.

Verifier study milestones

  1. Proposal and initial literature screening — published 6 October 2026.
  2. Full-text overlap review, benchmark manifest and baseline feasibility — next.
  3. Pilot, power analysis and frozen preregistration — planned.
  4. Confirmatory evaluation, technical report and independent reproduction — conditional on the earlier gates.

I welcome specific criticism of the protocol and collaborators interested in verification, code generation and reproducible inference research. Contact me about the study.