@gyro_ai: The most painful part of reproducing a machine learning paper is that the paper is vague, key parameters are hidden in the appendix or even not written at all, and you spend most of your time playing detective instead of writing code. paper2code is an Agent skill: give it an arxiv link, and it generates a runnable implementation code. 1308 stars htt…

X AI KOLs Timeline Tools

Summary

Paper2code is an AI Agent skill that generates runnable implementation code with citation anchors from an arxiv paper link, automatically audits ambiguities in the paper and marks unspecified parts, helping researchers and engineers efficiently reproduce machine learning papers.

Reproducing a machine learning paper is most painful when the paper is vague, key parameters are hidden in the appendix or even not written at all, and you spend most of your time playing detective instead of writing code. paper2code is an Agent skill: give it an arxiv link, and it generates a runnable implementation code. 1308 stars https://github.com/PrathamLearnsToCode/paper2code… MIT License. Input a paper link, output code traceable to its sources. Key features: 1. Every line of code is annotated with its source - The generated code marks which section and formula of the paper it corresponds to, e.g., "this line is written according to Section 3.2, Formula 4", making it easy to verify item by item. 2. Review ambiguities before writing - Before writing, each implementation choice is categorized into "clearly stated", "partially stated", or "not stated at all", unlike ordinary AI that silently fills in gaps. 3. Honestly mark uncertainties - Where the paper is unclear, the code directly marks it and lists several common approaches, rather than pretending to know. 4. Even dig into appendices - Appendices, footnotes, and figure captions are treated as serious sources, as these often contain key parameters. Install with one command: npx skills add. After adding to an Agent like Claude Code, type /paper2code plus a paper link to use it. You can also choose the framework and detail level. For graduate students and algorithm engineers who often need to reproduce papers, this can save you a lot of detective work.
Original Article
View Cached Full Text

Cached at: 05/23/26, 10:08 AM

Reproducing a machine learning paper is painful when the paper is vague—key parameters buried in appendices or omitted entirely, leaving you to play detective instead of writing code.

paper2code is an Agent skill. Give it an arxiv link, and it produces a runnable implementation. 1308 stars.

https://github.com/PrathamLearnsToCode/paper2code…

MIT license. Paper in, citation-anchored code out.

Core features:

  1. Line-by-line source annotation – Every line of generated code marks which section and equation of the paper it corresponds to (e.g., “this line follows section 3.2, Equation 4”), so you can verify each decision.
  2. Audit ambiguity before writing – Before generating any code, every implementation choice is classified as “specified by the paper,” “partially specified,” or “completely unspecified.” Unlike ordinary AI that silently fills in gaps.
  3. Honest uncertainty markers – Where the paper is unclear, the code directly marks the uncertainty with [UNSPECIFIED] tags and lists common alternatives, rather than pretending to know.
  4. Mine the appendices – Appendices, footnotes, figure captions are all treated as first-class sources—key parameters are often hidden there.

Install with a single npx skills add command. After installing into an agent like Claude Code, type /paper2code followed by the paper link to use it. You can also choose a framework and detail level.

For graduate students and algorithm engineers who frequently reproduce papers, this saves a lot of detective work.


PrathamLearnsToCode/paper2code

Source: https://github.com/PrathamLearnsToCode/paper2code

paper2code

arxiv URL in → citation-anchored implementation out

┌─────────────────────────────┐ ┌──────────────────────────────────────┐ │ │ │ {paper_slug}/ │ │ /paper2code │ │ ├── README.md │ │ https://arxiv.org/abs/ │ ───▶ │ ├── REPRODUCTION_NOTES.md │ │ 1706.03762 │ │ ├── requirements.txt │ │ │ │ ├── src/ │ │ │ │ │ ├── model.py # §3.2 cited │ │ │ │ │ ├── loss.py # §3.4 cited │ │ │ │ │ ├── train.py # §4.1 cited │ │ │ │ │ ├── data.py │ │ │ │ │ ├── evaluate.py │ │ │ │ │ └── utils.py │ │ │ │ ├── configs/ │ │ │ │ │ └── base.yaml # all params │ │ │ │ └── notebooks/ │ │ │ │ └── walkthrough.ipynb │ └─────────────────────────────┘ └──────────────────────────────────────┘

[placeholder: animated GIF showing the full pipeline — paper fetch → parsing → ambiguity audit → code generation → walkthrough notebook]


Why this exists

The problem: ML papers are vague. Critical hyperparameters are buried in appendices or omitted entirely. Prose contradicts equations. “Standard settings” refers to nothing specific. When you implement a paper, you spend more time detective-working than coding.

What LLMs get wrong: Naive code generation fills in every gap silently and confidently. You get something that runs but doesn’t match the paper. Worse, you can’t tell which parts are from the paper and which were invented by the model.

What paper2code does differently:

  1. Citation anchoring — every line of generated code references the exact paper section and equation it implements (§3.2, Eq. 4)
  2. Ambiguity auditing — before writing a single line of code, every implementation choice is classified as SPECIFIED, PARTIALLY_SPECIFIED, or UNSPECIFIED
  3. Honest uncertainty — unspecified choices are flagged with [UNSPECIFIED] comments at the exact line where the choice is made, with common alternatives listed
  4. Appendix mining — appendices, footnotes, and figure captions are treated as first-class sources, not ignored

The result: code you can trust because you can verify every decision against the paper.


Install

bash npx skills add PrathamLearnsToCode/paper2code/skills/paper2code

You’ll be prompted to:

  1. Select agents — pick the coding agents you want to use this skill with (e.g., Claude Code)
  2. Choose scope — Global (recommended) or project-level
  3. Choose method — Symlink (recommended) or copy

Once installed, open your agent and run the skill:

bash claude # or your preferred agent


Usage

Basic — generate a minimal implementation

/paper2code https://arxiv.org/abs/1706.03762

Specify framework

/paper2code https://arxiv.org/abs/2006.11239 --framework jax

Full mode — includes training loop and data pipeline

/paper2code 2106.09685 --mode full

Educational mode — extra comments and pedagogical notebook

/paper2code https://arxiv.org/abs/2010.11929 --mode educational

Using bare arxiv ID

/paper2code 1706.03762


What you get

attention_is_all_you_need/ ├── README.md # Paper summary, contribution statement, quick-start ├── REPRODUCTION_NOTES.md # Ambiguity audit, unspecified choices, known deviations ├── requirements.txt # Pinned dependencies ├── src/ │ ├── model.py # Architecture — every layer cited to paper section │ ├── loss.py # Loss functions with equation references │ ├── data.py # Dataset class skeleton with preprocessing TODOs │ ├── train.py # Training loop (if in scope) │ ├── evaluate.py # Metric computation code │ └── utils.py # Shared utilities (masking, positional encoding, etc.) ├── configs/ │ └── base.yaml # All hyperparams — each one cited or flagged [UNSPECIFIED] └── notebooks/ └── walkthrough.ipynb # Pedagogical notebook linking paper sections → code → sanity checks

Key files explained

FilePurpose
model.pyArchitecture only. Each class maps to a paper section. Variable names match paper notation.
REPRODUCTION_NOTES.mdThe ambiguity audit. Lists every choice, whether the paper specified it, and what alternatives exist.
base.yamlSingle source of truth for all hyperparameters.
walkthrough.ipynbRunnable on CPU with toy dimensions. Quotes paper passages, shows corresponding code, runs shape checks.

What this skill will NOT do

  • Won’t guarantee correctness. The implementation matches what the paper describes. If the paper is wrong, the code is wrong. If the paper is vague, the code flags it.
  • Won’t invent details. If the paper doesn’t specify a hyperparameter, the code uses a common default and marks it [UNSPECIFIED]. It will never silently fill in gaps.
  • Won’t download datasets. The data.py provides a Dataset class skeleton with clear instructions on where to get the data and how to preprocess it.
  • Won’t set up training infrastructure. No distributed training, no experiment tracking, no checkpointing beyond what the paper’s contribution requires.
  • Won’t implement baselines. Only the core contribution of the paper is implemented.
  • Won’t reimplement standard components. If the paper says “standard transformer encoder,” the code imports it or notes the dependency — it doesn’t reimplement attention from scratch.

Design principles

Citation anchoring convention

Every non-trivial code decision is anchored to the paper:

``python

§3.2 — “We apply layer normalization before each sub-layer” (Pre-LN variant)

class TransformerBlock(nn.Module): def forward(self, x): # §3.2, Eq. 2 — attention_weights = softmax(QK^T / sqrt(d_k)) attn_out = self.attention(self.norm1(x)) # (batch, seq_len, d_model) x = x + attn_out # §3.2 — residual connection ``

The UNSPECIFIED flag system

``python

[UNSPECIFIED] Paper does not state epsilon for LayerNorm — using 1e-6 (common default)

Alternatives: 1e-5 (PyTorch default), 1e-8 (some implementations)

self.norm = nn.LayerNorm(d_model, eps=1e-6) ``

``python

[ASSUMPTION] Using pre-norm based on “we found pre-norm more stable” in §4.1

The paper uses post-norm in Figure 1 but pre-norm in experiments — ambiguous

``

Ambiguity classification

TagMeaning
§X.YDirectly specified in paper section X.Y
§X.Y, Eq. NImplements equation N from section X.Y
[UNSPECIFIED]Paper does not state this — our choice with alternatives listed
[PARTIALLY_SPECIFIED]Paper mentions this but is ambiguous — quote included
[ASSUMPTION]Reasonable inference from paper context — reasoning explained
[FROM_OFFICIAL_CODE]Taken from the authors’ official implementation

Contributing

Adding worked examples

Worked examples are the most trust-building part of this project. To add one:

  1. Pick a well-known paper (people should be able to verify the output)
  2. Run the skill: /paper2code https://arxiv.org/abs/XXXX.XXXXX
  3. Save the full output to skills/paper2code/worked/{paper_slug}/
  4. Write a review.md that honestly evaluates:
    • What the skill got right
    • What it correctly flagged as unspecified
    • Any mistakes it made
    • Any edge cases it handled well or poorly
  5. Submit a PR with all of the above

Improving guardrails

If you find a pattern where the skill hallucinates or makes a silent assumption, add it to the appropriate file in guardrails/.

Adding domain knowledge

If papers in your subfield consistently reference components that the skill doesn’t know about (e.g., graph neural network primitives, RL components), add a knowledge file in knowledge/.


Worked examples

This repo includes fully worked examples to demonstrate output quality:

PaperTypeCommand
Attention Is All You Need (1706.03762)Architecture/paper2code https://arxiv.org/abs/1706.03762
DDPM (2006.11239)Training method/paper2code https://arxiv.org/abs/2006.11239

Each includes the complete generated output plus an honest review.md evaluating what the skill got right and wrong.

Similar Articles

@vintcessun: So this is how AI writing for academic papers can be done: not by direct polishing, but by first having the agent learn the target scenario, read excellent examples, and then record why each paragraph was written that way. PaperSpine uses a writing rationale matrix to turn writing into an auditable reasoning process, rather than black-box generation.…

X AI KOLs Timeline

PaperSpine is a paper writing skill suite for Codex, Claude Code, and OpenClaw, which transforms AI writing into an auditable reasoning process through a writing rationale matrix, rather than black-box generation.

@nini_incrypto_: Recommended Paper Writing Skills 1 Research-Paper-Writing-Skills https://github.com/Master-cai/Research-Paper-Writing-Skills… This is a skill pack for machine learning/computing...

X AI KOLs Timeline

Recommends four open-source paper writing skill packs suitable for machine learning/computer vision/NLP and other fields, focusing on structure standardization, polishing and review, complete research workflow, and Chinese collaboration, supporting AI assistants such as Codex, Claude Code, and Gemini.

@Xudong07452910: This might be the last paper written by humans for AI to read. Recently came across a paper co-authored by 37 authors from Stanford, CMU, Michigan, etc.: 'The Last Human-Written Paper'. The core point is quite bold: the centuries-old paper format may be outdated in the AI era...

X AI KOLs Timeline

A paper co-authored by 37 authors from Stanford, CMU, Michigan, etc. proposes ARA (Agent-native Research Artifact) to replace the traditional paper format, aiming to solve the narrative tax and engineering tax, enabling AI agents to understand, reproduce, and extend research.

@wsl8297: The biggest fear when writing a paper with AI is not being unable to produce content, but that it looks complete while the research gap, literature support, and argument structure are actually unsound. Academic Paper Skills solves this problem. GitHub: https://github.com/lishix520/…

X AI KOLs Timeline

Academic Paper Skills is a paper writing skill framework for Claude Code, dividing the writing process into two phases: strategy and composition. It incorporates features like literature support, reviewer simulation, and quality checks to help users generate a first draft from a research idea.