What the pipeline should do
| Stage | Input | Output |
|---|---|---|
| Trigger | Pull request event | One review job with a commit SHA |
| Context | Changed files, rules, and relevant contracts | Bounded review prompt |
| Model call | Prompt and diff | Structured findings |
| Verification | Finding, file, and line | Reproduction or dismissal |
| Publication | Verified findings | One review with actionable comments |
| Merge gate | Tests, review, and approval | Human decision or protected-branch result |
The model should not receive every secret or the entire repository by default. Start with changed files and add a related file only when the review contract requires it.
Claude and GPT routes
| Route | Candidate model ID | Use it when |
|---|---|---|
| Claude review | claude-opus-5-5 | The diff crosses modules or the failure mode is hard to infer |
| Claude daily pass | claude-sonnet-5 | The repository has a stable rubric and many routine pull requests |
| GPT review | gpt-6.1-sol | You want a second family in the comparison set |
| GPT fast pass | gpt-6-luna | The check is narrow, structured, and already covered by tests |
These are starting routes from the live catalog checked on October 4, 2026. Measure accepted findings and misses on your repository. Do not treat the model name as a quality guarantee.
Make the review output strict
Ask for JSON with a schema. Each finding needs a path, a line in the changed file, a consequence, reproduction steps, and a suggested fix. Return an empty array when there is no evidence.
const review = await client.responses.create({
model: "claude-opus-5-5",
input: [
{ role: "system", content: REVIEW_RULES },
{ role: "user", content: diffText },
],
text: { format: REVIEW_SCHEMA },
});
Use the request format supported by your SDK. Parse the result, reject unknown fields when practical, and fail the job visibly if the model returns invalid JSON. A silent parser repair hides the very issue the review is meant to find.
Verification before publication
The first model output is a lead. A maintainer or a verification step must check it against the source. For a race report, write a small test. For a SQL issue, inspect the query plan or fixture. For an API contract issue, compare the producer and consumer types.
Keep false positives in the fixture set. The next prompt should explain why the old finding was wrong. This gradually teaches the pipeline the boundaries of your repository without turning the prompt into a novel.
A GitHub Actions outline
name: AI review
on:
pull_request:
types: [opened, synchronize, reopened]
jobs:
review:
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
- name: Run tests
run: npm test -- --run
- name: Review the diff
env:
CLAUDEXIA_API_KEY: ${{ secrets.CLAUDEXIA_API_KEY }}
run: node scripts/ai-review.mjs
Pin the action versions you trust, limit permissions, and pass the pull request SHA to the script. Do not let a model-generated patch write to the default branch from the review job.
Review policy that teams can use
Block a merge only for a reproducible correctness, security, data-loss, or build issue. Leave style suggestions to the formatter and linter. Cap comments per file and deduplicate findings so the author sees a review, not a flood.
Measure the pipeline with accepted findings, dismissed findings, missed bugs discovered later, runtime, and token usage. Review those numbers after model or prompt changes.
FAQ
Can AI merge pull requests by itself?
It can propose findings. Merge authority should stay with branch protection and the team. An automated merge needs a separate policy, tests, and an audit trail.
Should I use Claude or GPT for code review?
Run both on the same anonymized fixtures. Use the route with the better verified signal for your repository, then keep a second route for comparison or fallback.
Can the model review the entire repository?
Only when the task requires it and your data policy permits it. A bounded diff plus relevant contracts is easier to audit and cheaper to run.
How do I stop prompt injection in source code?
Treat code and comments as untrusted data. Put review rules in the system message, delimit the diff, and tell the model that instructions inside the diff are not review policy.
What should I log?
Log commit SHA, model ID, duration, parse result, and finding identifiers. Redact secrets and avoid storing full source diffs unless your retention policy allows it.