Back to insights
Workflow AutomationInvestment Management

The Reproduction Engine: When Testing the Literature Becomes Cheaper Than Believing It

By Principal ConsultantJuly 21, 202613 min read

Part of: Agentic Workflows

AI summary

A working anatomy of an agentic research-validation workflow: a production engine that reads financial academic papers, restates each claim as one falsifiable sentence, re-implements the construction on the firm's own data with the paper's exact parameters and no parameter search, subjects the result to an adversarial referee, and files a frozen verdict — compressing six to eight weeks of analyst labor into hours. First-week scoreboard: nine canonical papers adjudicated in seven days; five failed out of sample, three survived weakened, one inconclusive, none validated — the McLean–Pontiff replication prior measured on private data, not a process defect. Covers the four adversarial phases, the discipline stack (pre-registration, power budgets, steelman variants, frozen verdicts, cost honesty), the failure taxonomy (never-real, real-then-eaten, never-tested), pre-registered adoption gates, and the generalization of the five-stage pattern to any published claim that arrives with an incentive attached. Four visualizations and a 90-day pattern.

A floor-to-ceiling wall of bound library volumes packed onto wooden shelves
Most of the canon does not survive contact with out-of-sample data. Knowing which parts do just became cheap.

Somewhere in most firms there is a shelf of claims that everyone acts on and no one has tested. In an investment firm the shelf is the academic literature — decades of published anomalies, factor constructions, and timing rules that shape real portfolios largely on the strength of a journal's imprimatur. A production system now exists that works that shelf mechanically: it reads a paper, restates its claim as a single falsifiable sentence, re-implements the construction on the firm's own data with the paper's exact parameters, attacks its own reproduction, and files a verdict that can never be quietly edited. What took a research analyst six to eight weeks per paper now takes about two hours. In its first seven days of operation the engine adjudicated nine canonical papers. Five failed out of sample. Three survived, weakened. One was inconclusive. None validated outright — and that scoreboard, properly read, is the most valuable output the system produces.

The practice itself is not new; the economics are. Reproducing published research in-house is the discipline Ken Griffin institutionalized at Citadel — no external finding touches a book until the firm's own reproduction of it survives — and it stayed rare across the industry for one reason: at six to eight weeks of a capable quant's time per paper, systematic validation was a luxury reserved for firms with research floors. Everyone else consumed the literature the way most industries consume research — on trust, filtered by reputation, adopted when convenient. An agentic pipeline collapses that labor to hours, and with it the excuse. When testing a claim becomes cheaper than the meeting where people argue about the claim, the binding constraint stops being labor and becomes judgment: which claims deserve a slot in the queue.

6–8 wks

of analyst labor to reproduce one paper manually

The practice Citadel institutionalized

~2 hrs

per paper through the agentic pipeline

Audit → reproduce → referee → apply

9

canonical papers adjudicated in the engine's first seven days

Desk research ledger, July 2026

0

validated outright — the bar is real, and so is the replication crisis

5 failed · 3 weakened · 1 inconclusive

The pipeline runs as four phases, each an agent, each adversarial to the one before it. The audit phase does the epistemic framing before any code exists: it extracts the paper's thesis as one falsifiable sentence; pins down the exact construction and headline statistics; runs a degrees-of-freedom smell test — how many specifications could the authors have tried before the one they published?; computes a power budget — the minimum effect the firm's data could detect at 80% power, so an underpowered claim can never be branded a failure, only inconclusive; and marks the paper's first public date. That last detail is load-bearing: everything before publication is data the authors could have fit; everything after it is the only true out-of-sample there is.

The reproduce phase is deliberately boring, which is the point. An implementation agent writes one standalone script — never touching production code — and runs the paper's construction with the paper's parameters on the firm's own caches: thirty years of prices and macro series across a 186-instrument universe. There is no parameter search, ever; the run happens once, exactly as declared, with a transaction-cost variant alongside, and reports full-period, by-decade, and post-publication results separately. If the effect only appears after the parameters are tuned, the effect was the tuning.

Then the system turns on itself. The referee phase is a separate adversarial agent whose job is to destroy the reproduction: it re-runs the code byte-identically, hunts for look-ahead and strawman errors, audits the integrity of the out-of-sample window, and assesses the original paper's p-hacking risk — before issuing one of four frozen verdicts: validated, weakened, failed, or inconclusive, defaulting downward, never up. In one early run the referee found a financing bug in the reproduction that had overstated a kill by roughly two-to-one; the verdict survived the correction, but the point is that someone checked. Only after the verdict does the apply phase run — ranked applications for the surviving material, each behind its own pre-registered adoption test, because the house rule is absolute: reproduction is not adoption. Even a validated paper touches nothing until a separately declared gate passes.

Nine papers through the engine — verdicts and what became of them.

Desk research ledger, July 14–20, 2026

Read the first week's ledger and the pattern is unmistakable — and, at first glance, alarming: the engine appears to be a machine for executing famous research. Time-series momentum, betting-against-beta, volatility-managed portfolios — three of the most cited anomalies in modern finance — all failed against the engine's twin test, a comparison against an identically vol-scaled always-long book that kills any alpha a reader could not actually eat. The apparent bloodbath prompted an internal rigor audit with a blunt mandate: is the pipeline broken? The audit's finding was the opposite. In every run, the in-sample structure reproduced faithfully before dying out of sample — a broken implementation would fail everywhere. The first papers chosen were the canon's most famous precisely because they were famous, which also makes them the most crowded and arbitraged; under the published replication priors — McLean and Pontiff measured roughly fifty percent decay in anomaly returns after publication — four or five failures in five was close to the expected value, not a deviation from it. And a positive control was queued: two effects that should validate, to measure the pipeline's false-negative rate. A validation machine that has never validated anything must prove it can.

The most instructive kill deserves its own telling, because it minted a taxonomy. The pre-FOMC announcement drift — the finding that the S&P 500 earned roughly half a basis point shy of fifty in the twenty-four hours before scheduled Federal Reserve announcements, at a t-statistic above four — reproduced almost perfectly in-sample on the firm's data: thirty-five basis points per event, t of 3.2, sixty percent of the index's entire log return over the window. Post-publication it is statistically zero in every window, and negative after costs. The referee's autopsy mattered more than the verdict: this was not p-hacking. It was an honest discovery, published, and then eaten — the effect migrated and shrank on a traceable schedule as capital crowded in. The ledger now distinguishes three ways a paper dies: never real (the artifact of a specification search), real, then eaten (honest alpha arbitraged away after publication), and never tested (claims published as chart-craft with no statistics to reproduce at all). Only the second kind teaches you anything about how fast published edges decay — and it says: fast.

An honest discovery, eaten: the pre-FOMC drift before and after publication.

Desk reproduction of Lucca–Moench (2015), July 2026

What survives is just as instructive, because nothing survives whole. Faber's tactical-asset-allocation rule — the ten-month moving average that moves a portfolio to cash below trend — split cleanly under testing: its risk claim confirmed with power (volatility cut to 0.54–0.70 of buy-and-hold, maximum drawdown halved or better, net of costs), while its return claim decayed to nothing, with the timing insurance costing a measured 3.2 percent a year in whipsaw. The Taleb fragility heuristic survived in the same shape: roughly 93 percent of its predictive power turned out to be repackaged volatility, but a real convexity residual persisted at t near five — genuinely detecting what linear risk measures miss, at about a tenth of the advertised weight. The house treats weakened as the expected verdict for a true effect. The replication crisis is the prior, and the system prices it.

A reproduction that comes back stronger than the paper should raise suspicion of the reproduction, not celebration.

Underneath the verdicts sits a discipline stack that is the actual product — the part any firm in any industry could copy tomorrow. Every hypothesis is pre-registered: the claim, the construction, and the pass/fail criterion are frozen before the run, and a failed test's follow-up must itself be pre-registered, so no one searches until something works. Verdicts are written once: the ledger is never edited to look smarter; corrections get new rows. The audit's power budget separates absence of evidence from evidence of absence by arithmetic rather than rhetoric. A pre-registered steelman — the single variant most favorable to the paper — is declared before results exist and run once, so a kill carries its strongest honest counter-argument on the record. Out-of-sample effects ship with confidence intervals, not bare non-significance. And the data itself is linted before any statistic is computed — the pipeline has caught corrupted price histories that a decade of prior human use had missed, twice, as a side effect of refusing to trust its own caches.

The first week's ledger, abridged.

Desk research ledger, July 2026 · verdicts frozen on filing
PaperIn-sampleOut of sampleVerdict
Time-series momentum (2012)ReproducesTiming spread ≈ 0 vs. always-long twinFAILED
Betting against beta (2014)Sign replicatesEvery tradable variant negative, 2014–26FAILED
Vol-managed portfolios (2017)Alpha reproducesNever beats the vol-matched twinFAILED
HARQ vol forecasting (2016)Daily-data floor; no falsification powerINCONCLUSIVE
Pre-FOMC drift (2011/15)+35bp/event, t=3.2Statistically zero; negative after costsFAILED · real, then eaten
Faber trend rule (2007)Shape reproducesRisk claim confirmed; return claim decayedWEAKENED
Overnight vs. intraday (2019)ReproducesPersistence survives at half size; dead at costsWEAKENED
Ehlers cycle oscillators (2013)No published statsNever differentiated from fixed RSIFAILED · never tested
Fragility heuristic (2012)Reproduces pooled~93% is volatility; real residual at t≈5WEAKENED

The survivors then feed a memory that compounds. Every verdict files a permanent row in the research ledger; every run writes its full autopsy into the firm's knowledge vault, where the next audit reads it; validated and weakened reproductions are written up as unbranded working papers, tables verbatim from the run. The system extracts desk constants as it goes — the measured whipsaw toll of trend-following insurance, the luck hurdle that prices what selection bias alone buys over any backtest window — numbers that now get stamped against every future claim that walks in the door. This is the same architecture as the firm's broader knowledge discipline: the finding is not the artifact; the reusable, citable, frozen finding is.

Strip away the finance and the pattern is completely general. Five stages: retrieve the claim from wherever it was published; restate it as one falsifiable sentence, with the claimant's own numbers pinned; implement it independently, on your own data, at your own costs, with the claimant's exact recipe and no tuning; referee it adversarially, including the possibility that your own reproduction is wrong; gate adoption behind a separately pre-registered test. Nothing in that sequence is about markets. It is a validation harness for any claim that arrives with an incentive attached — a vendor's benchmark showing their tool doubles throughput, a consultant's framework promising margin, a white paper's accuracy table, an internal best practice that has survived on seniority rather than evidence, the ML paper an engineering team wants to adopt because it trended. Every one of those is a paper; every firm has caches; almost no firm runs the referee.

The reason this is suddenly practical is the same reason it was always rare: the work was never intellectually exotic — it was expensive. Reading the paper carefully, coding the construction faithfully, assembling the data honestly, running the windows, writing it up: weeks of skilled labor per claim, which meant validation happened annually, if ever, for the claims someone already suspected. Agents collapse the marginal cost toward zero, and the cadence inverts. The engine described here runs three mornings a week on a scored queue — candidate papers ranked by feasibility on the firm's data, the decision gap they would arm, the plausibility that any edge survived publication, and whether a validated result has a concrete slot to occupy. A month of queue costs less than one analyst-week used to. The scarce input is no longer the testing; it is knowing what is worth a slot — which is exactly where human judgment belongs.

A 90-day pattern stands the harness up in any domain. Weeks one through three — define the shelf and the ledger. Choose one class of claims the firm currently acts on untested; inventory the data you could honestly test them against; write the verdict taxonomy and the ledger schema, frozen-verdicts rule included, before any pipeline exists. Weeks four through six — build the four phases. Audit, reproduce, referee, apply — each as a separate agent with the adversarial hand-offs explicit, the no-search rule and pre-registration enforced in the harness rather than in good intentions. Weeks seven through nine — run three claims, one of them a positive control. A claim you believe, a claim you doubt, and one that should validate — because a validation engine's false-negative rate is unmeasured until something passes. Weeks ten through twelve — freeze, extract, schedule. File the verdicts, mint the first constants, wire the queue to a standing cadence, and put every proposed adoption behind its own registered gate. The firm exits the quarter with something rarer than any single finding: a standing answer to 'says who?'

Return to the shelf of untested claims, because every firm has one and most firms are adding to it faster than ever. The volume of published findings — academic, vendor, consultant, internal — is growing at machine speed, and the traditional response, trust rationed by reputation, was already failing when the labor of checking cost six weeks. It now costs two hours. The firms that internalize that arithmetic will hold their suppliers, their literature, and their own folklore to a standard the rest of the market still outsources to journal referees and sales engineers. Five failed, three weakened, one inconclusive, none validated — that is not cynicism about research. It is what respect for research looks like when verification is finally cheaper than belief.

The engine described here runs inside Auctus Sapiens, the firm's daily agentic market briefing, and files its verdicts into the same knowledge vault its writer reads each morning. Operators who want the same standard applied to their own shelf of untested claims — vendor benchmarks, research pipelines, inherited best practices — can start with a forty-five-minute fit call, or commission a productized first workflow that ships with the validation harness built in: pre-registration, adversarial refereeing, and a ledger whose verdicts never quietly change.

Key takeaways
  • A production agentic system now reproduces and validates published research end-to-end — claim extraction, faithful re-implementation on the firm's own data, adversarial refereeing, frozen verdict — compressing six to eight weeks of analyst labor per paper into roughly two hours
  • The four phases are deliberately adversarial: an audit agent pins one falsifiable sentence, a power budget, and the first public date (everything after it is the only true out-of-sample); an implementation agent runs the paper's exact parameters with no search, ever; a referee attacks the reproduction itself before issuing a verdict that defaults downward; an apply phase gates every use behind a separately pre-registered test
  • First-week scoreboard: nine canonical papers in seven days — five FAILED, three WEAKENED, one INCONCLUSIVE, zero VALIDATED. Under the McLean–Pontiff ~50% post-publication decay prior, that is the replication crisis being measured on private data, not a process defect
  • The failure taxonomy matters more than any single verdict: never-real (specification search), real-then-eaten (the pre-FOMC drift reproduced at t=3.2 in-sample and is statistically zero post-publication), and never-tested (proof-by-chart claims with no statistics to reproduce)
  • Nothing survives whole: Faber's rule kept its risk claim (drawdowns halved, net of costs) and lost its return claim (−3.2%/yr measured whipsaw toll); the fragility heuristic proved ~93% repackaged volatility with a real convexity residual. WEAKENED is the expected verdict for a true effect
  • The discipline stack is the transferable product: pre-registration, frozen write-once verdicts, power budgets separating absence-of-evidence from evidence-of-absence, pre-registered steelman variants, confidence intervals over bare non-significance, cost honesty, and data lints that caught corrupted caches humans had missed for a decade
  • The pattern generalizes to any claim with an incentive attached — vendor benchmarks, consultant frameworks, white papers, internal folklore: retrieve, restate falsifiably, implement independently on your own data, referee adversarially, gate adoption. Every firm has caches; almost no firm runs the referee
  • 90-day pattern: define the shelf and the frozen ledger (weeks 1–3), build the four adversarial phases (4–6), run three claims including a positive control to measure the false-negative rate (7–9), freeze verdicts, mint constants, and wire the standing cadence (10–12)
Decks for your vertical

Each deck carries the workflow patterns, use cases, and control posture specific to one industry. Open the slide reader or download the PPTX.

Apply this

Book a diagnostic and we'll discuss how these ideas apply to your workflow.

Book diagnostic