Code in the Machine: Can Large Language Models Reproduce Econometric Analyses?

Aubrey Jolex*

Innovations for Poverty Action

January 28, 2026

Abstract

We test whether AI agents can replicate published econometric analyses from methodology text alone. Three models—GPT Codex, Claude Opus, and Gemini Pro—attempted to reproduce 13 J-PAL randomized trials using only methodology descriptions and raw data. Under standard prompts, execution rates varied (Codex 100%, Opus 85%, Gemini 77%) with volatile coefficient accuracy. Exhaustive variable-level specifications prompts closed the gap: all models achieved 100% execution, and Gemini achieved median drift of just 0.19 SE with 100% of estimates within 1.0 SE of published values. Divergences traced to prompt ambiguity (47%) and best-practice defaults (31%), not hallucinations (8%). We conclude that AI agents can aid replication when paired with structured prompts and validation protocols.

Keywords: AI Agents, Large Language Models, Reproducibility, Replication, Econometrics, Development Economics, Code Generation

JEL Codes: C18, C88, O12, A14

1. Introduction

The replication crisis in economics is well-documented. Chang and Li (2015) found that only 33% of economics papers could be replicated without author assistance, while Camerer et al. (2016) reported a 61% replication rate (significant effect in the same direction as the original study) for laboratory experiments. More recently, Brodeur et al. (2024) found that while 85% of papers were computationally reproducible, 25% contained coding errors and robustness checks reduced effect sizes by 52% on average. These findings raise fundamental questions about the reliability of published scientific research.

These persistent difficulties in replication stem in part from the complexity and opacity of modern empirical workflows. The growing use of large language models (LLMs) in empirical research may exacerbate, rather than resolve, these concerns. LLMs can now generate statistical code, translate methodological descriptions into executable analyses, and automate complex empirical workflows (Noy and Zhang, 2023; Korinek, 2025). By delegating core analytical tasks to probabilistic systems whose internal logic is not fully transparent, researchers risk introducing new sources of error that are difficult to detect using existing replication norms. When analytical pipelines are mediated by AI-generated code rather than fully human-authored programs, traditional notions of replication and transparency may no longer suffice. This shift raises a critical question: How can we verify and evaluate the reliability of econometric analyses produced with LLM assistance?

This question has direct implications for transparency and accountability in the age of AI. First, LLMs may lower the apparent cost of analysis while increasing the risk of undetected mistakes, particularly in the form of "silent logic errors"—code that executes without error but implements an incorrect model or statistic. Unlike conventional coding errors, which often require author assistance to diagnose, such errors may evade both peer review and standard replication attempts. Second, the widespread adoption of LLM-assisted workflows could scale these risks across the literature, potentially propagating systematic errors at a speed and scale that exceeds human-only research processes. Only under strict verification protocols could LLMs plausibly deliver offsetting benefits, such as lowering barriers to replication and enabling broader participation in research auditing. As Hamermesh (2007) observed, economists have traditionally treated replication as an ideal rather than a practice; without new safeguards, AI may further widen this gap rather than close it.

We address this question through a systematic audit of three frontier AI agents—GPT 5.1 Codex, Claude Opus 4.5, and Gemini 3 Pro, accessed via GitHub Copilot—attempting to replicate 13 randomized controlled trials (RCTs) from the J-PAL Dataverse. Our experimental design isolates AI capability from other factors: we provide standardized methodology prompts extracted from published papers, grant full access to replication datasets, but withhold original analysis code. This setup tests whether AI agents can implement a statistical analysis "from scratch" given only a description of methods.

Our findings are nuanced. Under initial methodology-level prompts, all three models achieved high computational reproducibility: GPT Codex completed 100% of papers, Claude Opus 85%, and Gemini Pro 77%. However, coefficient accuracy varied—Codex achieved near-perfect matches (within 0.05 SE) for 22% of outcomes, notably replicating the India Maternal Literacy RCT with coefficients of 0.0351 versus the published 0.035. An experimental manipulation with exhaustive variable-level prompts revealed two striking results: first, detailed specifications closed the execution gap entirely, with all three models achieving 100% success; second, coefficient accuracy improved dramatically, with Gemini achieving median drift of just 0.19 SE and 100% of estimates within 1.0 SE—outperforming both Codex and Opus under the same detailed conditions. Divergences traced primarily to prompt ambiguity (47%) and best-practice defaults (31%) rather than hallucinations (8%).

We make three contributions to transparency and verification in AI-assisted research. First, we provide the first systematic audit of AI agent replication capabilities in economics, establishing baselines for verification and accountability of AI-generated analyses. Testing multiple models on identical tasks reveals the "jagged frontier" of AI capability—areas where verification is essential versus where automation is reliable. Second, we develop a taxonomy of failure modes—hallucinations, misinterpretations, and best-practice deviations—that clarifies the explainability challenges specific to AI-generated econometric code. This taxonomy provides a framework for auditing AI-generated analyses and identifying when human oversight is critical. Third, we derive practical recommendations for methodology documentation standards that improve reproducibility for both human and machine replicators, addressing institutional and technical barriers to transparency in an era where AI tools are increasingly embedded in research workflows.

Our findings reveal that AI agents are neither silver bullets nor reckless saboteurs; their value depends on the structure embedded in prompts, the availability of validation checks, and researchers' willingness to interrogate generated code. In that respect, our exercise complements the productivity experiments of Noy and Zhang (2023) and the conceptual framing of Korinek (2023), but extends both by examining end-to-end replication of causal inference designs rather than discrete programming tasks. A key methodological insight emerges: reproducibility failures trace back to mundane implementation details such as naming conventions, scale transformations, and sample filters. AI agents amplify these issues because they dutifully implement whichever plausible interpretation the prompt permits, making structured documentation and runtime validation essential.

The stakes extend beyond academic housekeeping. As AI tools diffuse across universities, policy institutions, and consulting firms, the quality of prompts and validation protocols will determine whether automated replication becomes a democratizing force or a new source of specification drift. Our evidence suggests a path forward: structured documentation that makes identification logic explicit, runtime checks that catch implausible estimates before they circulate, and prompt-governance norms that treat instructions as first-class research objects. If the profession adopts these practices, AI agents can materially expand the scope and speed of credible empirical work. If not, we risk automating the garden of forking paths at scale.

The remainder of the paper proceeds as follows. Section 2 reviews the literature on replication failures, AI code generation, and transparency standards, situating our work within ongoing debates about automation and verification. Section 3 describes the 13 J-PAL randomized evaluations that comprise our benchmark, detailing selection criteria and key features. Section 4 outlines our experimental design, including prompt construction protocols, agent configurations, and evaluation metrics. Section 5 presents the main findings on computational reproducibility, coefficient accuracy, and failure modes. Section 6 discusses implications for research workflows, documentation standards, and AI governance, while Section 7 concludes with practical recommendations for integrating AI agents into replication and research pipelines.

2. Literature Review

The credibility literature has already established the stakes. Studies by Christensen and Miguel (2018), Brodeur et al. (2016), and Camerer et al. (2016) document publication bias, p-hacking, and low replication rates; Chang and Li (2015) famously reproduced only 51 percent of macro papers even with author help. More recent audits show limited progress despite data mandates (Herbert et al., 2021; Vilhuber, 2020), and the Institute for Replication's mass exercises (Brodeur et al., 2024) confirm that specification choices can halve effect sizes. If access to files is no longer the bottleneck, the remaining question is whether published methodology text contains enough detail for independent analysts—human or machine—to reconstruct the estimation logic.

This concern maps directly onto the "garden of forking paths" problem (Huntington-Klein et al., 2021). Even conscientious researchers make divergent decisions about variable construction, clustering, or sample restrictions. Pre-analysis plans help but raise concerns about rigidity (Coffman and Niederle, 2015), while sensitivity frameworks such as Cinelli and Hazlett (2020) show how fragile many estimates remain. AI-generated code could become a middle ground: it can expose undocumented forks quickly, yet it also risks multiplying silent mistakes if prompts lack precision.

Economists increasingly view AI as a "third co-author" (Korinek, 2023). Experiments show that language models already mimic economic reasoning (Horton, 2023), boost productivity (Noy and Zhang, 2023), and operate along a jagged capability frontier (Dell'Acqua et al., 2023). Software engineers warn that the main danger is not syntax errors but silent logic errors (IEEE, 2025), motivating tools such as RepoAudit (2025) and data-cleaning agents (Zhang et al., 2025). These insights imply that econometric workflows must emphasize prompt clarity and runtime validation rather than blind trust in fluent code.

Our paper contributes to this conversation in three ways. First, we revisit the replication crisis literature through the lens of automation. Classic literature such as Hamermesh (2007) framed replication as a public good problem with weak rewards and high opportunity costs. We show that LLM agents change the cost curve but do not remove the informational asymmetry: unless methodology sections include machine-readable detail, agents simply formalize the same forks that human replicators confront. Second, we connect the transparency movement (Christensen and Miguel, 2018; Vilhuber, 2020) to concrete tooling requirements. The reason transparency policies have plateaued is that "available" files are not synonymous with "understandable" logic. Our prompt treatments operationalize the level of structure needed for both humans and machines to succeed, thereby extending the data-availability conversation toward instruction-availability.

Third, we bridge the AI governance literature with practical econometric needs. Spirling et al. (2025) argue that stochastic language-model outputs complicate traditional definitions of replication, yet their remedy emphasizes process documentation. Our findings echo that prescription: the most reliable way to harness LLMs is to log prompts, runtime checks, and resolution steps so that downstream users can audit the full chain of reasoning. At the same time, the jagged frontier documented by Dell'Acqua et al. (2023) predicts that capabilities will improve unevenly across tasks. We therefore catalogue which elements of development economics replications (e.g., clustered OLS) already cross the automation threshold and which (e.g., subgroup ANCOVA or IV) still require bespoke oversight.

Figure 1: Positioning our contribution within AI code-generation research

Simple Syntax
Human-in-Loop
AI assistance
Complex Causal Logic
Human-in-Loop
Traditional replication
Simple Syntax
Fully Autonomous
Benchmarks (HumanEval)
Complex Causal Logic
Fully Autonomous
This paper

Notes: Prior work studies either simple syntax tasks or human-in-the-loop causal inference. We benchmark fully autonomous causal logic.

3. The Research Sample: 13 Randomized Evaluations

We selected thirteen J-PAL randomized evaluations because they mirror the everyday workload of development economists: multi-arm interventions, clustered standard errors, attrition adjustments, and occasional nonlinear models. The portfolio covers public health (COVID-19 behavior campaigns in China, HIV curricula in Kenya and Cameroon, Zambian water purification pricing), private-sector interventions (Mexican SME consulting, South African credit enforcement), governance reforms (Indonesian corruption audits, Pakistani tax incentives), and education programs (Indian maternal literacy, Canadian achievement awards). Sample sizes range from 245 loans to 150,000 patient encounters, and every study ships with publicly accessible replication packages.

Rather than reproduce each narrative in detail, we highlight the features most relevant for testing AI agents. Some designs depend on simple OLS with clustered errors (Indonesia corruption case study), others require binary or count models (Kenya HSV-2 prevention, U.S. physician messaging), and several hinge on intricate sample filters or subgroup splits (Canadian second-year males, maternal literacy ANCOVA with baseline controls). Appendix A lists the full study roster. The breadth of contexts ensures that our benchmark probes the same specification decisions—outcome choice, fixed effects, clustering, and transformations—that routinely derail human replications.

Two practical considerations shaped the final paper selection. First, each replication archive had to be fully self-contained so that an AI agent operating inside a VS Code workspace could ingest the raw data without manual preprocessing. We therefore excluded otherwise-relevant studies whose replication materials depended on proprietary modules, missing dictionaries, or bespoke operating-system configurations. Second, we prioritized studies with rich methodological prose. Papers that simply restated regression equations without describing variable construction offer little fodder for prompt engineering; conversely, narratives that spelled out treatment definitions, balance checks, and attrition rules allowed us to craft prompts that mimic a careful reader's notes.

4. Experimental Design and Methods

4.1 Study selection and preprocessing

We filtered the J-PAL Dataverse for randomized evaluations that met four practical constraints: open replication packages, standard econometric estimators (OLS/logit/Poisson rather than bespoke structural models), methodology text rich enough to build prompts, and complete data files. For each study we inspected the package, documented file relationships, and ensured that any replication failure would reflect AI capability rather than missing inputs.

Preprocessing followed a standardized and repeatable workflow. All archives were cloned into a uniform directory structure, and file integrity was verified against the corresponding Dataverse records. When archives contained multiple study arms or auxiliary appendices, readme stubs were created to map filenames to the tables described in the publications. These procedures reflect the minimal preparation a human replicator would typically undertake prior to coding, thereby ensuring that the experiment isolates the translation from prose to estimation rather than extensive forensic data cleaning.

4.2 Prompt construction

Baseline prompts paraphrased the published methodology: research question, treatment definitions, target tables, outcome descriptions, control sets, fixed effects, sample restrictions, and clustering levels. We intentionally mirrored real papers by leaving vagueness where authors did. No original code snippets were provided. The second detailed-prompt treatment rewrote the same instructions with exhaustive specificity—verbatim variable names, missing-data flags, transformation rules, and logical operators for every filter—so we could observe whether "more detail" alone solved replication gaps.

To minimize experimenter degrees of freedom, we drafted prompt templates before examining any model output. Each template contains four blocks: (i) a high-level goal statement, (ii) data management instructions, (iii) estimation requirements, and (iv) reporting expectations. Within each block we substituted study-specific text while preserving the scaffolding. The exhaustive prompts added a fifth block containing literal variable dictionaries and sample logic expressed in pseudo-code. This design allows us to attribute performance differences to the content of the instructions rather than to ad hoc stylistic changes.

4.3 Agent configuration

All experiments ran through GitHub Copilot's agent routing system via Visual Studio Code. Each of the three models—GPT 5.1 Codex, Claude Opus 4.5, and Gemini 3 Pro—was tested in a dedicated workspace to ensure independence and reproducibility. Agents received identical prompts, had access to Stata 18.5 or R 4.5 depending on the replication archive, and could iterate after runtime errors. We logged iteration counts, execution time, and final code. Agents worked independently; no model saw another's output. Behavioral differences emerged organically: Codex favored extensive diagnostics, Claude prioritized concise scripts, and Gemini persevered through repeated debugging.

4.4 Evaluation and diagnostics

Replication quality is measured via coefficient drift, the absolute difference between AI and published coefficients divided by the published standard error. We classify drift into perfect (<0.05 SE), minor (0.05–0.20), moderate (0.20–0.50), substantial (0.50–1.0), and major (≥1.0) categories to separate cosmetic differences from consequential deviations. After every divergent estimate we compared AI code to author scripts (when available) and assigned root causes to four buckets: hallucination, misinterpretation, best-practice default, or data mismatch. These diagnostics underpin the failure taxonomy reported later.

5. Results

5.1 Detailed prompts close the execution gap

Exhaustive instructions were not a silver bullet. Table 1 reveals a striking non-monotonic relationship: providing variable-level specifications improved Opus and Gemini but collapsed Codex's performance entirely.

Table 1: Initial vs. Detailed Prompt Performance
Model Execution Success Perfect Matches
Initial Detailed Initial Detailed
GPT 5.1 Codex 13/13 (100%) 13/13 (100%) 8/36 (22%) 1/5 (20%)
Claude Opus 4.5 11/13 (85%) 13/13 (100%) 2/44 (5%) 1/6 (17%)
Gemini 3 Pro 10/13 (77%) 13/13 (100%) 0/17 (0%) 5/7 (71%)

Notes: Execution success = papers that produced executable code and regression output. Perfect matches defined as coefficient drift <0.5 SE from published values. Detailed prompts achieved universal execution: all three models reached 100%, with Opus improving by 15 percentage points and Gemini by 23 points. Notably, Gemini's coefficient accuracy improved dramatically under detailed prompts (71% within 0.5 SE, 100% within 1.0 SE), achieving the best accuracy of all three models.

Codex maintained its 100% execution rate under both prompt conditions, demonstrating robust performance regardless of specification detail. This consistency suggests that Codex's architecture handles both concise methodology descriptions and exhaustive variable-level specifications equally well.

Opus's improvement—from 85% to 100% execution—reveals that its binding constraint was prompt ambiguity rather than econometric capability. Under initial prompts, Opus failed on Papers 02 (Mexico SME) and 13 (Pakistan tax) due to unclear variable mappings and sample restrictions. The detailed prompts eliminated this ambiguity, allowing Opus to proceed systematically through each requirement.

Gemini's substantial improvement—from 77% to 100%—traces directly to variable-naming disambiguation. Under initial prompts, Gemini persistently struggled with data dictionary mappings, attempting to infer outcome variables from imperfectly labeled datasets and frequently selecting the wrong column. Explicit variable names in detailed prompts solved this mechanical problem across all papers. This 23 percentage point gain—the largest among the three models—suggests that Gemini's binding constraint was purely informational rather than architectural.

The key finding is that detailed prompts achieved universal execution success. All three models reached 100% under detailed specifications, completely eliminating the execution gap observed under initial prompts. This implies that explicit variable-level documentation is sufficient to overcome the ambiguity-driven failures that plagued initial attempts. The practical implication is clear: invest in detailed methodology documentation, as it enables even the weakest performer to match the strongest. Figure 2 visualizes these patterns.

Figure 2: Execution Success by Prompt Condition

100%
100%

GPT Codex
±0pp

85%
100%

Claude Opus
+15pp

77%
100%

Gemini Pro
+23pp

Initial Prompts Detailed Prompts

Notes: Execution success rates under initial prompts (standard methodology descriptions) versus detailed prompts (exhaustive variable-level specifications). Annotations show percentage point change. All three models achieved 100% execution under detailed prompts: Codex maintained its lead, Opus improved by 15 percentage points, and Gemini improved by 23 points. N = 13 papers per model-condition pair.

5.2 Overall replication success

Table 2: Model Performance Summary
Model Papers Completed Success Rate Coefficients Generated Perfect Match Time (min)
GPT 5.1 Codex 13/13 100% 431 8 (22%) ~180
Claude Opus 4.5 11/13 85% ~44 2 (5%) ~95
Gemini 3 Pro 10/13 77% 17 0 (0%) ~240

Notes: Papers Completed indicates the number of papers for which the model generated executable code producing coefficient estimates. Perfect Match defined as drift < 0.05 SE. Time includes all iterations and debugging.

Codex is the only agent that combined universal execution with double-digit perfect matches, albeit at the cost of long runtimes. Claude worked roughly twice as fast but sacrificed accuracy and failed on two multilevel designs. Gemini persevered through repeated debugging but rarely landed on the correct specification. In short: Codex for precision, Claude for quick exploratory replications, Gemini for stubborn data wrangling.

5.3 India maternal literacy: a success example

The maternal literacy and CHAMP experiment highlights how well agents perform once prompts specify variables precisely. Codex essentially reproduced the main table:

Table 3: India Maternal Literacy Study: Coefficient Comparison
Outcome Published β Codex β Drift (SE) Match
Child Math (ML) 0.035 0.0351 0.01 Perfect
Child Math (CHAMP) 0.032 0.0324 0.03 Perfect
Child Math (ML+CHAMP) 0.056 0.0557 0.02 Perfect
Child Language (ML+CHAMP) 0.042 0.0423 0.02 Perfect
Mother Total (ML+CHAMP) 0.12 0.1237 0.28 Minor

Notes: Coefficients represent standard deviation units. Drift calculated as the absolute difference between AI-generated and published coefficients divided by published standard errors. ML = Maternal Literacy; CHAMP = Children's homework support program.

Codex succeeded because it matched the author recipe almost verbatim: normalized outcomes, the full baseline-control vector with missing indicators, stratum fixed effects, and village clustering. Claude, by contrast, dropped the controls and triggered collinearity—underscoring that small prompt omissions map directly into major estimation gaps.

5.4 Coefficient Accuracy Under Detailed Prompts

Table 4: Coefficient Accuracy Under Detailed Prompts
Metric GPT Codex Claude Opus Gemini Pro
Coefficients matched 5 6 7
Mean drift (SE) 2.42 1.74 0.32
Median drift (SE) 0.93 0.84 0.19
Within 0.5 SE 20% 17% 71%
Within 1.0 SE 60% 67% 100%
Within 2.0 SE 80% 83% 100%

Notes: Drift calculated as |β̂LLM − βpub| / SEpub. Manual matching based on alignment between LLM output variable names and published outcome descriptions. Sign concordance across all matches: 94% (17/18).

Under detailed prompts, Gemini achieved the best coefficient accuracy: mean drift of 0.32 SE (median 0.19 SE), with 71% of estimates within 0.5 SE and 100% within 1.0 SE. This represents a remarkable improvement from its poor initial performance and demonstrates that detailed prompts can dramatically improve coefficient accuracy for models that struggled with ambiguity. Figure 3 shows the distribution of coefficient drift under initial prompts.

Figure 3: Distribution of Coefficient Drift by Model (Initial Prompts)

Percentage of Outcomes
70%
Perfect
<0.05
Minor
0.05-0.20
Moderate
0.20-0.50
Substantial
0.50-1.00
Major
≥1.00
GPT Codex Claude Opus Gemini Pro

Drift Category (in Standard Errors)

Notes: Drift categories defined as: Perfect (<0.05 SE), Minor (0.05–0.20 SE), Moderate (0.20–0.50 SE), Substantial (0.50–1.00 SE), Major (≥1.00 SE). Percentages based on all coefficients generated by each model under initial methodology-level prompts. GPT Codex shows the best distribution with 22% perfect matches, while Gemini Pro shows 64% major deviations under initial prompts (improving dramatically under detailed prompts).

5.5 Divergence Source Classification

Table 5: Divergence Source Classification
Source Frequency Example
Outcome variable selection 34% Papers 01, 05, 12: multiple outcomes available
Sample restriction missed 23% Papers 01, 12: subgroup specifications
Scale/transformation 19% Papers 04, 06: proportions vs. percentage points
Control specification 15% Papers 09, 11: baseline controls omitted
Data encoding issues 9% Papers 05, 08: variable name mismatches

Notes: Classification based on systematic comparison of AI-generated code to original analysis scripts. Percentages sum to 100% of identified divergence sources.

5.6 Hallucinations are rare; ambiguity is not

Only 8 percent of divergences involved outright hallucinations in which a model invented a method that contradicted the prompt. For the most part, agents respected econometric guardrails.

The remaining gaps were mundane: 47 percent misinterpretations where vague prose allowed multiple defensible implementations, and 31 percent best-practice defaults such as reporting proportions instead of percentage points in the Indonesia audit study. These patterns mean analysts must read the generated code, rescale units, and decide whether the difference reflects presentation or substance—the same diligence already required for human-written replications.

6. Root Cause Analysis

6.1 Taxonomy of Failure Modes

Table 6: Failure Mode Taxonomy
Category Definition Responsibility Frequency
Hallucination AI agent invented method not in prompt AI Error 8%
Misinterpretation Ambiguous text led to wrong choice Prompt Gap 47%
Best Practice Applied standard-but-unspecified approach Methodological 31%
Data Mismatch Variable names differed from prompt Documentation 14%

Notes: Classification based on systematic code comparison for all divergent results. "Responsibility" indicates primary attribution for divergence.

6.2 The Jagged Frontier in Econometrics

Consistent with Dell'Acqua et al. (2023), we find AI agent capability is uneven across tasks:

Table 7: Econometric Task Difficulty for AI Agents
Capability Level Technique Success Rate Common Failures
High Basic OLS 85% Scale issues
Clustered SE 78% Wrong clustering level
Summary statistics 95% Minor formatting
Moderate Fixed effects 72% Wrong FE specification
ANCOVA with controls 55% Missing baseline vars
Multi-arm RCT 45% Treatment coding
Low Subgroup analysis 35% Sample restrictions
IV/2SLS 30% Instrument selection
Complex DiD 40% Timing, parallel trends

Notes: Success rate indicates percentage of attempts producing coefficients within 1 SE of published values. Based on combined performance across all three models.

The table makes the frontier obvious: agents cruise through plain OLS, summary stats, and even clustered errors, but accuracy falls off once specifications demand precise baseline controls, multi-arm logic, or sample restrictions. Anything that requires translating prose labels into exact variable lists—ancova controls, subgroup filters, IV instruments—remains fragile.

7. Discussion

7.1 Cross-Cutting Themes: Transparency, Standards, Open Science

Our evidence speaks to three intertwined fronts that matter across journals, funders, and policy labs: transparency in the age of AI, adaptive standards that balance privacy with reproducibility, and open-science practices that build trust.

Transparency and accountability for AI systems. By benchmarking multiple agents on public RCTs and publishing the full prompt-plus-validation loop, we demonstrate how to verify generative technologies with concrete drift metrics and failure taxonomies. The runtime critique protocol functions as an accountability layer that any lab can adopt when deploying AI in predictive or generative settings.

Adapting standards and tooling. The prompt treatments, documentation templates, and checklists provide a blueprint for encoding identification logic without exposing sensitive data. They highlight where privacy-preserving variable descriptions suffice and where literal column names are indispensable, helping practitioners navigate the trade-off between confidentiality and replicability.

Open science, trust, and reuse. Version-controlled prompt logs, automated plausibility checks, and structured replication archives lower the marginal cost of auditing results. By showing that simple guardrails catch every catastrophic failure, we offer evidence that transparency investments can shift perceptions among policymakers and spread the benefits of AI assistance more evenly across research teams.

7.2 Practitioner's Checklist: Evidence-Based Recommendations

Table 8: Practitioner's Checklist for AI-Assisted Replication
Stage Action Evidence from This Study
Before Running AI Agent
Prompt Design Detailed prompts eliminate execution failures All three models achieved 100% under detailed prompts (Opus +15pp, Gemini +23pp)
Variable Names Include exact column names, not prose descriptions Papers with SD units (explicit scaling) had 57% lower drift (0.84 vs 1.94 SE)
Complexity Assessment Classify paper difficulty; allocate more review time for hard papers Easy papers averaged 0.77 SE drift; hard papers averaged 3.34 SE drift
During Execution
Sample Size Check Verify N matches published value within 5% Sample mismatch explained 14% of failures
Iteration Monitoring Flag runs exceeding 5 iterations as high-risk Failed Codex runs averaged 6.1 iterations vs 3.4 for successes
After Execution
Magnitude Check Flag coefficients >2 SE from expectations 100% of catastrophic failures would have been caught by 2-SE threshold
Sign Verification Confirm coefficient sign matches theory Sign errors were present in 15% of major divergences
Clustering Audit Verify clustering level matches paper Wrong clustering was most common best-practice failure

Notes: Recommendations derived from analysis of 39 agent-paper pairs across 13 J-PAL RCTs. SE = standard error of published coefficient.

7.3 Limitations

Four caveats remain. (i) The thirteen RCTs intentionally use mainstream estimators; we do not test more exotic designs such as weak-IV settings, RDD bandwidth searches, or structural simulation. (ii) Prompt wording reflects our interpretation of each methodology section. Authors writing their own structured prompts might supply richer context or embed new biases. (iii) Models evolve quickly, so our benchmark is a snapshot of early-2025 capabilities, not a permanent ranking. (iv) Each agent-paper pair was run once; output variance documented by Spirling et al. (2025) implies that repeated draws could widen or narrow the drift distribution.

8. Conclusion

Our benchmark shows that frontier agents easily write executable code but only reproduce coefficients when methodology text supplies the right level of detail. Under initial prompts mimicking standard paper descriptions, Codex executed 100% of studies (matching 22% of coefficients exactly), Opus 85%, and Gemini 77%. Providing exhaustive variable-level specifications closed the execution gap entirely: all three models achieved 100% success, with Opus improving by 15 percentage points and Gemini by 23 points. This universal success demonstrates that detailed documentation eliminates ambiguity-driven failures regardless of model architecture. Overall divergences mostly trace back to prompt ambiguity (47%) and best-practice defaults (31%) rather than hallucinations (8%), and simple plausibility checks would have caught every catastrophic miss.

The takeaway is to invest less in ever-longer prompts and more in structured documentation plus runtime validation. Prompt templates that force analysts to name variables, filters, and clustering units, combined with automated checks on magnitudes and sample sizes, can harness AI speed without sacrificing credibility. With those guardrails, agents become practical collaborators for replication and, by extension, for new empirical work.

Looking ahead, three avenues appear most promising. First, integrating self-auditing routines—similar to the validation loop we piloted—directly into IDEs would allow AI agents to flag inconsistent sample sizes or implausible magnitudes before analysts even inspect the output. Second, training datasets that pair methodology text with executable code (for example, curated replication packages annotated with prompts) could teach models to infer missing details more reliably. Third, journals and data repositories can incorporate prompt logs as part of submission materials so that future researchers inherit not only scripts but also the instructions used to generate them. Collectively, these improvements would shift the conversation from "Can AI replicate?" to "How do we design institutions so that AI and humans jointly safeguard credibility?"

Acknowledgements

We thank the Abdul Latif Jameel Poverty Action Lab (J-PAL) for making replication data publicly available through their Dataverse repository. We declare no conflicts of interest.

References

Brodeur, A., Lé, M., Sangnier, M., & Zylberberg, Y. (2016). Star Wars: The Empirics Strike Back. American Economic Journal: Applied Economics, 8(1), 1–32. https://doi.org/10.1257/app.20150044

Brodeur, A., Mikola, D., Cook, N., et al. (2024). Mass reproducibility and replicability: A new hope. IZA Discussion Paper No. 16912. https://www.iza.org/publications/dp/16912

Camerer, C. F., Dreber, A., Forsell, E., et al. (2016). Evaluating replicability of laboratory experiments in economics. Science, 351(6280), 1433–1436. https://doi.org/10.1126/science.aaf0918

Chang, A. C., & Li, P. (2015). Is economics research replicable? Sixty published papers from thirteen journals say "usually not." Finance and Economics Discussion Series 2015-083, Board of Governors of the Federal Reserve System. http://dx.doi.org/10.17016/FEDS.2015.083

Christensen, G. S., & Miguel, E. (2018). Transparency, reproducibility, and the credibility of economics research. Journal of Economic Literature, 56(3), 920–980. https://doi.org/10.1257/jel.20171350

Cinelli, C., & Hazlett, C. (2020). Making sense of sensitivity: Extending omitted variable bias. Journal of the Royal Statistical Society: Series B, 82(1), 39–67. https://doi.org/10.1111/rssb.12348

Coffman, L. C., & Niederle, M. (2015). Pre-analysis plans have limited upside, especially where replications are feasible. Journal of Economic Perspectives, 29(3), 81–98. https://doi.org/10.1257/jep.29.3.81

Dell'Acqua, F., McFowland III, E., Mollick, E. R., et al. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School Working Paper 24-013. https://www.hbs.edu/faculty/Pages/item.aspx?num=64700

Hamermesh, D. S. (2007). Replication in economics. NBER Working Paper No. 13026. https://doi.org/10.3386/w13026

Herbert, S., Kingi, H., Stanchi, F., & Vilhuber, L. (2021). The reproducibility of economics research: A case study. Banque de France Working Paper No. 853. https://publications.banque-france.fr/en/reproducibility-economics-research-case-study

Horton, J. J. (2023). Large language models as simulated economic agents: What can we learn from homo silicus? NBER Working Paper No. 31122. https://doi.org/10.3386/w31122

Hu, C., Zhang, L., Lim, Y., Wadhwani, A., Peters, A., & Kang, D. (2025). REPRO-BENCH: Can agentic AI systems assess the reproducibility of social science research? Findings of the Association for Computational Linguistics: ACL 2025. https://aclanthology.org/2025.findings-acl.1210

Huntington-Klein, N., Arenas, A., Beam, E., et al. (2021). The influence of hidden researcher decisions in applied microeconomics. Economic Inquiry, 59(3), 944–960. https://doi.org/10.1111/ecin.12992

IEEE Transactions on Software Engineering. (2025). Advancing LLM-generated code reliability: A hybrid approach for hallucination detection. IEEE Transactions on Software Engineering.

Korinek, A. (2023). Generative AI for economic research: Use cases and implications for economists. Journal of Economic Literature, 61(4), 1281–1317. https://doi.org/10.1257/jel.20231736

Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586

RepoAudit. (2025). An autonomous LLM-agent for repository-level code auditing. arXiv:2501.18160. https://arxiv.org/abs/2501.18160

Spirling, A., Barrie, C., & Palmer, A. (2025). Replication for language models: Problems, principles, and best practices for political science. Working paper, November 2025.

Vilhuber, L. (2020). Reproducibility and replicability in economics. Harvard Data Science Review, 2(4). https://doi.org/10.1162/99608f92.4f6b9e67

Zhang, J., Khan, J., Zenil, H., et al. (2025). Exploring the role of large language models in the scientific method. Nature Communications.


* Corresponding author: aubreyjolex@gmail.com
This working paper is for discussion purposes. Please do not cite without permission.