How to Set Up Codex for Legacy Code Migrations

OpenAI Codex is a lightweight coding agent that runs in your terminal, reads a project-level AGENTS.md, and chains tool calls across your repo until a goal is met. Most teams use it for greenfield work. The leverage is in the opposite direction.
The thesis of this piece is simple. The single best use of Codex in 2026 is brownfield legacy migration: Python 2 to 3, jQuery to React, AngularJS to Angular, monoliths to services, untested codebases to tested ones. Per EffectiveSoft's modernization survey, AI-assisted migrations now run 2–3 times faster than manual rewrites with full test parity. Greenfield Codex is fun. Brownfield Codex pays the bills.
Setup matters more here than anywhere else. A Codex install pointed at a legacy repo with a vague AGENTS.md will hallucinate confidently, "modernize" working code, and ship subtle bugs that survive review. Done correctly, the same install ports thousands of lines a week with passing tests.

The Quick Answer
What is Codex for legacy migration?
A terminal agent that reads your repo, follows rules in AGENTS.md, and uses subagents to plan, edit, and validate code changes in batches small enough for humans to review. It is neither autocomplete nor a one-shot rewriter. It is a junior engineer who never tires and never improvises.
Why does it matter?
Legacy migration is the highest-friction work in software, and the friction is exactly the kind agents handle well: repetitive transformations, surface-level pattern matching, mechanical refactors with verifiable test gates. Per Lanil Marasinghe's React 19 migration write-up, one team migrated 20,000+ unit tests in weeks, not months.
The non-obvious truth: Codex does not replace the senior engineer who designs the migration. It replaces the four junior engineers who would otherwise type it.
The Foundations
Codex installs as a CLI: npm install -g @openai/codex, then run codex in any project directory. Three pieces of context wire it up.
The first is AGENTS.md, a markdown file at your repo root. Per the official guide, this file is concatenated directly into the context window at the start of every session, meaning whatever you put in AGENTS.md is the first thing Codex reads before it sees any prompt. For a legacy migration this is your project's constitution: source language, target language, framework versions, build commands, test commands, allowed and forbidden patterns, the migration plan itself.
The second is subagents. Per OpenAI's subagents docs, Codex can spawn specialized agents in parallel and collect their results in one response. For migration this is the difference between sequentially porting 200 files and porting 20 files at once with a coordinator agent gating the merge.
The third is skills. Skills are folders with a required SKILL.md and optional supporting files that bundle a procedure Codex can run on demand. ThePromptBuddy's sister piece on how to set up Claude Code for real SEO work covers the same primitive on Anthropic's stack. The shape is the same. The vendor and the failure modes differ.
The pattern is convergent across vendors. Cursor 3 isn't an IDE anymore, it's a control room, and Codex from the terminal does the same job from a different surface. The teams that win are not the ones picking sides. They are the ones who understand the primitive.
Why is legacy migration the killer use case for Codex?
Three properties of legacy work map exactly onto what agents do well. The work is mechanical, the success criteria are testable, and the value is high enough that even partial automation pays the bill.
Mechanical. Most legacy transformations are not creative. Replace componentWillMount with useEffect. Convert print x to print(x). Swap $.ajax calls with fetch. Rename module imports across 800 files. Codex with the right AGENTS.md handles these the same way every time. Per LegacyLeap's jQuery modernization data, agentic AI tools combined with the React Compiler now automate roughly 40% of code refactoring while maintaining functional parity. The other 60% is where humans still earn their keep.
Testable. The defining feature of brownfield work is that the existing system already has correct behavior, even if the code is ugly. That gives you the most valuable thing in agent workflows: a verifiable success gate. Tests pass or they don't. Output matches the reference or it doesn't. Behavior preserved or it isn't. Codex agents that can be evaluated automatically can be deployed at scale without supervision drift.
High-value enough to pay for itself. Greenfield AI work is consumer surplus. Brownfield AI work is enterprise budget. A 2026 React migration project that quotes $1.2M and 14 months turns into $300K and 4 months when an agent ports tests and components in batches. The math is the math. Engineering managers with budget read content like this for the budget code, not the demo.
The deeper reason this is the killer use case is timing. Frontier-model reasoning crossed the threshold somewhere in late 2025 where multi-file refactors with cross-cutting concerns finally land cleanly. Before that, agents on legacy code were a liability. After that, they are a competitive moat. The teams shipping legacy migrations on Codex right now are buying time-to-market that competitors who wait another year will not catch up to.
How should you actually structure AGENTS.md for a brownfield repo?
The document should answer five questions in order, in plain prose, with concrete code examples, in roughly 200–600 lines depending on project size.
1. What is this codebase?
Three sentences. Source language and version. Target language and version. The business that depends on it. Codex will hallucinate context if you do not provide it. Concrete beats abstract: "Django 1.11 + Python 2.7, migrating to Django 4.2 + Python 3.12. Powers checkout for a B2B marketplace serving 80M monthly transactions."
2. What are the migration rules?
A bullet list of allowed transformations and forbidden ones. Forbid is more important than allow. Forbidden transformations are where unsupervised agents drift. Examples that earn their place:
- "Never rewrite SQL queries during file ports. Migration is a separate phase."
- "Never change function signatures. If a signature must change, surface a question, do not auto-edit."
- "Never modify files in
legacy/payments/. Surface for human review."
3. How do you verify a change?
Exact commands, in order. pytest path/to/file -xvs, then mypy path/to/file --strict, then pre-commit run --files <changed>. Codex agents are remarkably good at obeying explicit verification chains. They are remarkably bad at inventing them.
4. What are the known sharp edges?
A numbered list of the five to ten gotchas your senior engineers complain about. The undocumented coupling between two services. The middleware that runs in a non-obvious order. The fixture that mutates global state. Codex cannot discover these on its own, but it will respect them when documented. This is where domain knowledge becomes leverage.
5. What does done look like?
A specific exit criterion per phase. Not "all tests pass." That is true on day one. Specifically: "Phase 2 done when tests/payments/ passes under Python 3.12 with no from __future__ imports remaining."
A common failure mode is treating AGENTS.md as a prose preamble. It is not. It is a configuration file with prose syntax. The teams getting the most out of Codex on migration projects revise their AGENTS.md weekly as the migration reveals new sharp edges, and they treat it as a first-class artifact in version control.

When does Codex break down on legacy code (and what to do)?
Three places where the abstraction leaks. The same three places where senior engineering judgment becomes non-negotiable.
Implicit business logic in untested code paths. Legacy systems often encode business rules in side-effects, accidental coupling, and "we don't touch that file because it broke last time." Codex cannot recover rules that exist only in operational tribal knowledge. The fix is not better prompting. The fix is investing in characterization tests before migration begins. Anthropic's zero-day work showed the same pattern: Claude finding bugs that survived 27 years of human review, but only when wired with the right tooling and supervised by researchers who knew what to look for. Substitute "researchers" with "domain engineers," and the lesson holds for migration.
Cross-cutting concerns at scale. Codex handles single-file ports cleanly. It struggles when a single change must propagate through 40 files with subtle variation. The mitigation is to script the cross-cutting transform manually as a codemod, then point Codex at the cleanup of edge cases the codemod missed. Tools handle the mechanical 90 percent. Codex handles the messy 10 percent.
Dependency upgrades disguised as migrations. A "Python 2 to 3" project is rarely just a language port. It is also a Django 1.11 to 4.2 leap, a celery 3 to 5 migration, and a Postgres 9.6 to 14 cutover. Codex will happily attempt all four at once and produce changes that compile and ship sideways into production incidents. The fix is sequencing in AGENTS.md: explicit phase gates, one dependency at a time, with characterization tests between each phase.
The model you pick matters too. Claude Opus 4.7 vs GPT-5.5 comes down to four questions, and for legacy migration the right answer is usually GPT-5.5 inside Codex on long agentic tasks because it owns the integrated tool surface and ships with the AGENTS.md primitive baked in. Hard-coding one model for every task type is the same setup mistake teams make on greenfield agents and pay double for on legacy.
The Edge Cases and Breakages
Generated code and vendored dependencies. Codex cannot tell the difference between vendor/, node_modules/, generated protobuf stubs, and your hand-written code unless AGENTS.md tells it. Add an explicit do-not-edit section listing every directory the agent must skip. Otherwise expect a session to "improve" autogenerated code and waste a half day.
Test gaps masked as test passes. A migration that flips green tests to green tests means nothing if the original tests had 12 percent coverage. Codex will not flag this. Run a coverage diff between source and target as part of the success criteria, and treat any decrease as a regression even when tests pass.
Cost discipline. Long agentic sessions on large codebases burn tokens. A site-wide brownfield migration that runs unsupervised for a week can run $5K to $20K in API costs depending on model selection and subagent depth. Cap session budgets, enforce per-batch limits, and log token spend per phase. The default mode of "let it run" is how a $200 weekend tool becomes a $5K monthly tool by accident.
The Bottom Line
Codex is the right tool for legacy migration in 2026 because the work is mechanical, testable, and high-value, and an agent in a terminal is the only shape that handles all three at once.
Set it up with a precise AGENTS.md that names the source and target stacks, lists allowed and forbidden transformations, encodes the verification chain, documents the sharp edges, and specifies the exit criteria per phase. Use subagents for cross-cutting passes. Sequence dependency upgrades. Cap session budgets. Treat the agent as a tireless junior engineer, not as a senior architect with a code editor.
Then ignore every guide that tells you to one-shot-rewrite a 400K-line monolith over a weekend. That guide is a Twitter demo, the codebase will not be migrated, and the team that tried will be doing characterization tests for the next six months anyway.
FAQ
Is Codex free?
The CLI is free to install. You pay for API usage through your OpenAI account, billed by token. A solo migration project on a small repo can run productively on $100 to $300 a month. An enterprise team running parallel subagents on a 400K-line codebase should budget $2K to $20K monthly depending on phase intensity.
Can Codex replace a migration consultancy?
For tactical execution, mostly yes. For migration strategy and phase sequencing, no. Most agencies that survive 2026 will be smaller teams using Codex on the implementation, not larger teams ignoring it on principle.
Will Codex break my legacy code?
It will, the first week, until your AGENTS.md catches up. The fix is to start on a non-production branch, run characterization tests on every change, and accept that the first phase of any migration is partly an AGENTS.md tuning exercise. Per the official Codex CLI features page, sandboxing and approval modes exist for exactly this.
Which migrations are best suited to Codex?
Language version bumps (Python 2 to 3, Java 8 to 17, Node 14 to 22), framework leaps with strong codemod support (AngularJS to Angular, jQuery to React), test framework migrations (Mocha to Jest, JUnit 4 to 5), and type adoption projects (untyped JS to TypeScript). Pure rewrites of business logic into a different paradigm are still mostly senior-engineer work.
Do I need to learn the OpenAI Agents SDK to use Codex?
For ad-hoc migration work, no. The CLI plus an AGENTS.md is enough. For repeatable workflows you run across many repos, the Agents SDK integration is worth the learning curve.
How is Codex different from Claude Code or Cursor for migration?
Claude Code is a terminal agent on the Anthropic stack with the same primitive shape. Cursor is an IDE-first orchestration surface where the agent lives inside the editor. For migration specifically, the terminal-first agents (Codex, Claude Code) tend to win because legacy work is more about repo-wide passes than file-by-file editing. Most teams running serious migrations end up with one terminal agent for batch ports and one IDE agent for hand-edits.
Is it safe to give Codex write access to a production repo?
Only with hooks, sandboxing, and a staging branch in place. Never wire production write access on day one. Start with read-only mode, build trust through plan-only sessions, and expand permissions phase by phase.