AI-Assisted Merge-Request Review
The flagship use case: a CI job that reads a merge request's diff, asks a model to review it, and posts the feedback back onto the MR — automatically, on every merge request. You'll build it end to end, with the model call shown in both Claude and OpenAI, and learn why review is the task where the reasoning-heavy model earns its cost.
Learning objectives
- Capture a merge request's diff inside a CI job.
- Send the diff to a model with a review-focused prompt (Claude or OpenAI).
- Post the model's review back as an MR comment via the GitLab API.
- Wire it all into
.gitlab-ci.ymlas a merge-request-only job.
.gitlab-ci.yml and set the secrets (gl5 covers doing that safely).1 · The shape of an AI reviewer essential
An automated reviewer is three moves: get the diff, ask the model, post the reply. Everything else is plumbing. The diff is what changed in the MR; the model is a senior engineer you're renting by the token; the posted comment is how the feedback reaches the author where they're already looking. None of these moves is novel on its own — you've made model calls since the API chapter — the newness is only that a push triggers it and the result goes back to GitLab.
Why automate review at all, when a human will review anyway? Because the AI reviewer never gets tired, never skips the boring files, and runs in seconds on every MR — so it catches the obvious stuff (missing error handling, a hardcoded secret, an off-by-one, an undocumented public function) before a human spends attention on it. It doesn't replace human review; it does the first, tireless pass so humans review the reviewed.
The common mistake is pointing the model at the whole repository instead of the diff. Reviewing everything on every MR is slow, expensive, and noisy — the author changed 40 lines, not 40,000. Scope the review to the diff (plus just enough surrounding context), and both cost and signal improve dramatically.
- The blue box is the input — the MR diff, not the whole repo (that scoping is the key efficiency).
- The middle box is the model call — the one piece shown in both Claude and OpenAI below.
- The green box is the output — the review posted back to the MR via GitLab's API, where the author sees it.
In short: capture → review → post. The model call is familiar; the CI trigger and the post-back are what make it a reviewer.
2 · The review prompt & the model call intermediate
The heart of the reviewer is a function that takes the diff and returns review text. The prompt frames the model as a focused reviewer — concrete, prioritized, and honest about severity. Here it is in both SDKs; the review logic is identical, only the vendor call differs.
The review call — in Claude or OpenAI
The system prompt and the diff handling are identical; only the SDK call changes. Review is a reasoning-heavy task, so this is where you'd point the more capable model (and dial up effort). Toggle the tab to switch SDK.
review.pyimport sys
from anthropic import Anthropic
client = Anthropic()
SYSTEM = (
"You are a senior code reviewer. Review the diff only. "
"List concrete issues by severity (blocker/major/minor), each with "
"file:line and a fix. If it's clean, say so briefly. No praise padding."
)
def review(diff: str) -> str:
resp = client.messages.create(
model="claude-opus-4-8", max_tokens=1500,
system=SYSTEM,
messages=[{"role":"user","content": diff}],
)
return next(b.text for b in resp.content if b.type=="text")
if __name__ == "__main__":
diff = open(sys.argv[1]).read()
print(review(diff))
review.pyimport sys
from openai import OpenAI
client = OpenAI()
SYSTEM = (
"You are a senior code reviewer. Review the diff only. "
"List concrete issues by severity (blocker/major/minor), each with "
"file:line and a fix. If it's clean, say so briefly. No praise padding."
)
def review(diff: str) -> str:
resp = client.responses.create(
model="gpt-5.5", max_output_tokens=1500,
reasoning={"effort": "high"}, # review is reasoning-heavy
instructions=SYSTEM,
input=diff,
)
return resp.output_text
if __name__ == "__main__":
diff = open(sys.argv[1]).read()
print(review(diff))
- The
SYSTEMprompt is the whole quality lever: it tells the model to review only the diff, prioritize by severity, citefile:line, propose fixes, and skip empty praise. - The diff comes in as a command-line argument (the file the CI job wrote in gl1) and is passed as the user content /
input. - Only the vendor call differs —
client.messages.create→ content blocks on Claude,client.responses.create→output_texton OpenAI (with the effort dial turned up, since review rewards reasoning).
Try this: tighten the prompt to your team's standards — "flag any public function without a docstring," "block on any hardcoded credential." The reviewer is only as good as the rubric you give it.
3 · Posting the review back to the MR intermediate
A review nobody sees is useless. GitLab has a REST API for posting a note (comment) to a merge request; the CI job calls it with the review text, authenticating with a token and using the predefined MR variables from gl1 to target the right MR.
post_comment.pyimport os, sys, urllib.request, urllib.parse
def post_note(body: str):
project = os.environ["CI_PROJECT_ID"]
mr_iid = os.environ["CI_MERGE_REQUEST_IID"]
token = os.environ["GITLAB_TOKEN"] # a CI secret (gl5)
url = f"https://gitlab.com/api/v4/projects/{project}/merge_requests/{mr_iid}/notes"
data = urllib.parse.urlencode({"body": body}).encode()
req = urllib.request.Request(url, data=data, method="POST",
headers={"PRIVATE-TOKEN": token})
urllib.request.urlopen(req)
if __name__ == "__main__":
post_note(sys.stdin.read()) # review text piped in
- GitLab's notes endpoint (
/merge_requests/:iid/notes) posts a comment; theCI_PROJECT_IDandCI_MERGE_REQUEST_IIDvariables (from gl1) target the exact MR. - Auth is a
PRIVATE-TOKENheader — a project/group access token stored as a masked CI variable (never in the file; gl5 covers secrets). - The review text is piped in on stdin, so the pipeline step is just
python review.py changes.diff | python post_comment.py.
Try this: GitLab also supports inline diff comments (on specific lines) via the discussions API — a more advanced reviewer parses the model's file:line references and posts them inline instead of as one summary note.
4 · Wiring it into the pipeline advanced
Now assemble it: a merge-request-only job that captures the diff, reviews it, and posts the result. This is the gl1 skeleton with the two scripts dropped in.
.gitlab-ci.ymlai-review:
stage: test
image: python:3.12-slim
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
script:
- pip install anthropic openai # whichever vendor review.py uses
- git fetch origin $CI_MERGE_REQUEST_TARGET_BRANCH_NAME
- git diff origin/$CI_MERGE_REQUEST_TARGET_BRANCH_NAME > changes.diff
- python review.py changes.diff | python post_comment.py
allow_failure: true # advisory — don't block merge on reviewer errors
- The job runs only on MRs (
rules:), on a Python image, installing the SDK yourreview.pyuses. - It captures the diff, then pipes
review.py's output straight intopost_comment.py— the three moves from section 1, as one script line. allow_failure: truemakes the reviewer advisory: a model timeout or API hiccup posts nothing but never blocks the merge. (gl5 discusses when to make an AI check blocking vs advisory.)
Try this: set ANTHROPIC_API_KEY or OPENAI_API_KEY as a masked CI variable and open a test MR — the review appears as a comment within a minute. That round trip is the whole feature.
5 · Tech-lead — making the reviewer trusted tech-lead
An AI reviewer that cries wolf gets muted. The lead-level work is making it trusted: keep it advisory at first (allow_failure: true) so a flaky model never blocks a merge; tune the prompt to your real standards so it flags what matters and stays quiet otherwise; and label its comments clearly as automated so humans weight them appropriately. Measure whether its findings get acted on — if developers routinely dismiss them, the rubric is wrong, not the developers. A reviewer earns the right to become a required check only after it's demonstrably precise.
🪜 Practice ladder beginner → industry
- Beginner: run
review.pylocally against a saved diff file and read the output. - Easy: tune the
SYSTEMprompt to add one team-specific rule. - Core: wire the MR-only job (Lab G2.3) and open a test MR to see the comment.
- Stretch: switch the review from Claude to OpenAI via the tabs and compare the feedback.
- Hard: parse the model's
file:linerefs and post inline diff comments via the discussions API. - Industry: add a metric tracking how often AI findings are acted on, and decide advisory-vs-required from it.
✓ Checkpoint — you can move on when you can…
- Capture an MR diff in a CI job.
- Review it with a model (Claude or OpenAI) using a severity-focused prompt.
- Post the review back as an MR comment via the GitLab API.
- Wire the whole thing as an advisory, MR-only job.
Knowledge check check yourself
Why scope an AI reviewer to the MR diff rather than the whole repository, and why keep the job allow_failure: true at first?
Show answer
allow_failure: true makes the reviewer advisory: a model timeout or API error posts nothing but never blocks the merge, which is how you introduce an AI check without it becoming a flaky gate before it's proven precise.What are the three moves of the AI reviewer, and how does the review text get back to the author?
Show answer
POST /projects/:id/merge_requests/:iid/notes), authenticated with a masked token and targeted using the predefined CI_PROJECT_ID and CI_MERGE_REQUEST_IID variables, so the comment lands on the exact MR where the author is working.