Skip to content
Claudexia TeamCODE REVIEW

AI code review pipeline with Claude or GPT: a safe CI design

Build a Claude or GPT code review pipeline that reads diffs, returns structured findings, runs tests, and leaves merge authority with the team.

What the pipeline should do

StageInputOutput
TriggerPull request eventOne review job with a commit SHA
ContextChanged files, rules, and relevant contractsBounded review prompt
Model callPrompt and diffStructured findings
VerificationFinding, file, and lineReproduction or dismissal
PublicationVerified findingsOne review with actionable comments
Merge gateTests, review, and approvalHuman decision or protected-branch result

The model should not receive every secret or the entire repository by default. Start with changed files and add a related file only when the review contract requires it.

Claude and GPT routes

RouteCandidate model IDUse it when
Claude reviewclaude-opus-5-5The diff crosses modules or the failure mode is hard to infer
Claude daily passclaude-sonnet-5The repository has a stable rubric and many routine pull requests
GPT reviewgpt-6.1-solYou want a second family in the comparison set
GPT fast passgpt-6-lunaThe check is narrow, structured, and already covered by tests

These are starting routes from the live catalog checked on October 4, 2026. Measure accepted findings and misses on your repository. Do not treat the model name as a quality guarantee.

Make the review output strict

Ask for JSON with a schema. Each finding needs a path, a line in the changed file, a consequence, reproduction steps, and a suggested fix. Return an empty array when there is no evidence.

const review = await client.responses.create({
  model: "claude-opus-5-5",
  input: [
    { role: "system", content: REVIEW_RULES },
    { role: "user", content: diffText },
  ],
  text: { format: REVIEW_SCHEMA },
});

Use the request format supported by your SDK. Parse the result, reject unknown fields when practical, and fail the job visibly if the model returns invalid JSON. A silent parser repair hides the very issue the review is meant to find.

Verification before publication

The first model output is a lead. A maintainer or a verification step must check it against the source. For a race report, write a small test. For a SQL issue, inspect the query plan or fixture. For an API contract issue, compare the producer and consumer types.

Keep false positives in the fixture set. The next prompt should explain why the old finding was wrong. This gradually teaches the pipeline the boundaries of your repository without turning the prompt into a novel.

A GitHub Actions outline

name: AI review
on:
  pull_request:
    types: [opened, synchronize, reopened]

jobs:
  review:
    permissions:
      contents: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v4
      - name: Run tests
        run: npm test -- --run
      - name: Review the diff
        env:
          CLAUDEXIA_API_KEY: ${{ secrets.CLAUDEXIA_API_KEY }}
        run: node scripts/ai-review.mjs

Pin the action versions you trust, limit permissions, and pass the pull request SHA to the script. Do not let a model-generated patch write to the default branch from the review job.

Review policy that teams can use

Block a merge only for a reproducible correctness, security, data-loss, or build issue. Leave style suggestions to the formatter and linter. Cap comments per file and deduplicate findings so the author sees a review, not a flood.

Measure the pipeline with accepted findings, dismissed findings, missed bugs discovered later, runtime, and token usage. Review those numbers after model or prompt changes.

FAQ

Can AI merge pull requests by itself?

It can propose findings. Merge authority should stay with branch protection and the team. An automated merge needs a separate policy, tests, and an audit trail.

Should I use Claude or GPT for code review?

Run both on the same anonymized fixtures. Use the route with the better verified signal for your repository, then keep a second route for comparison or fallback.

Can the model review the entire repository?

Only when the task requires it and your data policy permits it. A bounded diff plus relevant contracts is easier to audit and cheaper to run.

How do I stop prompt injection in source code?

Treat code and comments as untrusted data. Put review rules in the system message, delimit the diff, and tell the model that instructions inside the diff are not review policy.

What should I log?

Log commit SHA, model ID, duration, parse result, and finding identifiers. Redact secrets and avoid storing full source diffs unless your retention policy allows it.