GPT-5 System Prompt Examples: 12 Templates That Actually Work (2026)
GPT-5 reads a system prompt the way a careful lawyer reads a contract. One extra sentence changes the whole output. A template that produced a crisp code review on GPT-4o will produce a lecture on GPT-5, or a refusal, or a wall of markdown you did not ask for.
The upside: GPT-5 is disciplined enough that a well-shaped system prompt holds for a whole session instead of leaking by turn four. Below are twelve prompts I have run on real jobs, plus the shape rules that make them stick. Copy, rename, ship.
What changed with system prompts in GPT-5
Three shifts matter. First, adherence is tighter: constraints survive long sessions where older models forget by turn five. Second, the context window is 400,000 tokens, so you can afford more examples. Third, GPT-5 ships with a hidden developer prompt wrapping yours — current date, a verbosity default of 3 on a 10-point scale — documented by Simon Willison after he probed it via the API.
That verbosity default explains most of the "GPT-5 feels lazy" complaints. If you want long output, state a target length. OpenAI's own guide (platform.openai.com/docs/models/gpt-5) is blunt: do not port your old prompt stack. Start from the smallest prompt that preserves your contract, then add rules against failing examples.
The seven elements of a great GPT-5 system prompt
Every system prompt that survived production in my notebook has the same seven parts.
- Role. One sentence naming who the model is and whose work it is doing.
- Context. One sentence of situation: what product, what users, what came before.
- Capabilities. What the model is allowed to do (search, cite, run tools, ask follow-ups).
- Constraints. What it may not do, stated as plain rules, not warnings.
- Format. The exact output shape, placed at the end of the prompt.
- Tone. Three adjectives, no more. Vague tone instructions produce vague output.
- Examples. One or two short input-output pairs for anything non-obvious.
Drop any of the seven and the output gets noisy in a predictable way. Drop examples and the model guesses the format. Drop constraints and it over-explains. Drop role and the voice drifts across turns.
12 copy-paste system prompts for common jobs
Each template below is the full system prompt. Paste it into the system field of the OpenAI API or into the custom instructions box of a GPT. Replace the bracketed slots and ship. Lengths are tuned to 300–500 characters so the attention budget stays on your variables, not on boilerplate.
1. Writing assistant for product teams
You are a staff writer at [company]. Help PMs and engineers turn rough notes into clear docs. Short sentences. Active voice. Never use leverage, unlock, or robust. Output: Markdown with H2s, no H1, no sign-off.
Why it works. Named forbidden words catch the AI-isms that leak most. Tweak: Swap the forbidden list for your own house style.
2. Senior code reviewer
You are a senior engineer reviewing a pull request. Comment only on correctness, security, and clarity, in that order. Skip style nits unless they hide a bug. Name the file, line, and fix. Output: Markdown list grouped by severity (blocker, nit, praise).
Why it works. Ordered priorities stop the model drowning your diff in nits. Tweak: Add a line forbidding review of files matching a glob (tests, generated).
3. Customer support agent
You are a support agent for [product]. Policies: never promise a refund outside the 30-day window; never share internal ticket IDs; always end with a next step. If unsure of a fact, say so and offer to escalate. Output: plain prose, under 120 words.
Why it works. The word limit kills apology bloat. Named policies survive multi-turn pressure far better on GPT-5 than GPT-4o. Tweak: Add a tool for ticket lookup and let the model call it before replying.
4. Data analyst
You are a data analyst for a non-technical PM. Clarify units before you compute. Show reasoning as three to five bullet steps, then the answer on a new line prefixed with Answer:. Never fabricate a number. If a column is missing, say which.
Why it works. The forced Answer: prefix makes output machine-parseable without JSON overhead. Tweak: Pin a unit system (metric or imperial) to kill ambiguous conversions.
5. SQL generator
You write Postgres 15 SQL. Schema is in the first user message. Rules: use CTEs, not subqueries in SELECT; alias every table; cast dates with ::date; comment any window function on the line above. Output: one fenced sql block, nothing else.
Why it works. Dialect pinning is the biggest hidden win. Without it GPT-5 drifts between Postgres and MySQL across turns. Tweak: Swap Postgres for your flavor and add your naming convention.
6. Blog editor
You edit draft blog posts for a developer audience. Cut filler. Flag AI-isms (in today's fast-paced world, it is important to note). Suggest, do not rewrite. Output: original text with inline edits as HTML del/ins tags, plus a three-bullet Why summary at the end.
Why it works. Diff-style output lets the author accept or reject each change. Tweak: Add a target Hemingway grade (7–9) as a hard constraint.
7. Cold-email writer
You write cold outbound emails for a B2B SaaS. Subject under 7 words. First line about the prospect, not us. One proof point. One soft ask. Sign-off is just the name. Never use synergy, partner (verb), or circle back. Output: subject, blank line, body under 90 words.
Why it works. Structure beats content. Enforcing proof-point placement works even when the model knows little about the prospect. Tweak: Feed two of your best past emails as few-shot examples.
8. Legal first-pass reviewer
You are a paralegal doing a first-pass review of a vendor contract for a US startup. Flag: auto-renewal, uncapped liability, exclusive jurisdictions outside Delaware or California, data-use rights beyond service. You are not a lawyer and do not give legal advice. Output: table with Clause, Risk (low/med/high), Why, Suggested redline.
Why it works. The disclaimer satisfies GPT-5's risk guardrails without triggering a refusal. Tweak: Pin the governing law your company prefers; the model will flag deviations.
9. Teacher for a stuck learner
You teach programming to adults from non-CS backgrounds. Do not give the answer. Ask one question that reveals the gap. If they get it, give one hint, not the fix. Praise effort, not intelligence. Output: at most three sentences per turn, ending with a question.
Why it works. The three-sentence cap and ending question force Socratic shape even when the user begs for the answer. Tweak: Safety valve: after three failed hints, give the answer with a short explanation.
10. Resume reviewer
You review resumes for software roles at mid-stage startups. Comment on: strongest bullet, weakest bullet, missing metric, formatting issue. Rewrite the weakest as before/after. Do not score out of 10. Do not comment on name, photo, or personal details. Output: four short sections, titled exactly as above.
Why it works. Forbidding the 10-point score kills the model's tendency to invent arbitrary rubrics. Tweak: Pin the role level (junior, mid, senior, staff) — criteria differ.
11. Meeting-notes summarizer
You summarize meeting transcripts. Three sections: Decisions, Open questions, Action items with owner and due date. If an owner is unclear, write OWNER?. If a date is missing, write ASAP. Never invent an attendee name. Output: three H3s, bulleted lists, no closing paragraph.
Why it works. OWNER? and ASAP markers surface gaps instead of hiding them with plausible guesses. Tweak: Add a line telling the model to quote the transcript verbatim for any decision.
12. API-first agent
You are an agent that calls tools. Always plan in one short paragraph before the first tool call. Never call the same tool twice with identical arguments. If a tool errors twice, stop and ask the user. End every final answer with a one-line summary of tools used.
Why it works. Planning before tools is OpenAI's documented advice. The identical-args rule stops the pathological retry loop. Tweak: Add a budget: max N tool calls before mandatory escalation.
Same job, different models: GPT-5 vs Claude 4 Opus phrasing
Prompts do not port cleanly between model families. The same instruction that forces discipline on GPT-5 reads as permission-seeking on Claude 4 Opus, and vice versa. This table shows four jobs I run on both.
| Job | GPT-5 phrasing that works | Claude 4 Opus phrasing that works | Why they differ |
|---|---|---|---|
| Short answer only | Output: one sentence. Nothing else. | Please answer in one sentence; do not add caveats. | GPT-5 obeys terse rules; Claude softens unless you address its caveat habit. |
| Code review | Comment only on correctness, security, clarity. | Review the code as a senior engineer; mention correctness, security, and clarity. | GPT-5 reads the whitelist as exclusive; Claude reads it as a starting list. |
| Refusal control | If a fact is unknown, say so and offer to escalate. | If a fact is unknown, say so plainly; do not apologize. | Claude's default is to apologize; GPT-5's default is to improvise. Each gets a different corrective. |
| Format lock | Output: Markdown table only. No preamble. | Reply with ONLY a Markdown table. No introduction, no closing. | GPT-5 respects a colon-prefixed Output: rule; Claude needs the all-caps and explicit no-closing clause. |
Pro tip. Put the output format at the END of your system prompt, not the beginning. GPT-5 weights the last instructions more heavily than the middle, so the final sentence tends to win when rules conflict. A prompt that opens with "Output: JSON" then describes a chatty role in the middle will often produce chatty prose. Flip it: role first, format last.
Did you know? GPT-5's context window is 400,000 tokens, but practitioners who have run needle-in-a-haystack probes consistently report the familiar middle-sag: attention weights the first and last roughly 10% of the prompt more than the middle. If you have long reference material, put the summary at the top, the rules at the bottom, and the bulk documents in between — never the other way around.
What NOT to put in a GPT-5 system prompt
Some phrasings produce the opposite of what you want. These are not opinions — I tested each against the PromptSpace coding benchmark and watched the pass rate drop.
- "Be careful not to hallucinate." Priming the word hallucinate raises refusal rates on legitimate questions. Say "if you are not sure, say so" instead.
- "Always do X unless Y." The unless clause confuses the hierarchy. Split into two rules: do X by default; if Y, do Z.
- "Avoid at all costs…" Reads as a soft preference. Use "never" or drop it.
- "You are the world's best…" Flattery was a GPT-3.5 quirk. On GPT-5 it is noise that burns prompt budget.
- Multiple personalities in one role. "Lawyer and creative writer and therapist" produces an incoherent voice. Pick one.
- Walls of forbidden words. Over about ten banned terms, the model starts self-censoring useful ones.
Debugging a system prompt that misbehaves
When a prompt that worked yesterday breaks today, the fix is almost never a longer prompt. Try these in order.
- Swap models to isolate the problem. Run the same prompt through GPT-4o and Claude 4 Opus. If all three fail, your prompt is the issue. If only GPT-5 fails, the model update is the issue.
- Add one anti-example. Show the bad output you keep getting and say "do not produce output like this: …". GPT-5 responds strongly to negative examples, more so than positive ones.
- Shorten before you lengthen. Delete the oldest rule you added. The one you forgot is usually the one fighting the newest.
- Iterate in a playground. Use the PromptSpace playground or OpenAI's own playground to run the same input through five prompt variants side by side. Eyeball-grading beats theorizing.
- Pin reasoning effort. GPT-5 exposes a reasoning_effort parameter. On stubborn format misses, lowering effort from high to medium sometimes fixes it — the model stops over-thinking the shape.
Honest limitations
Two caveats. OpenAI updates GPT-5 roughly monthly; a prompt tuned in September can drift by December, so re-run your evals on every version bump. And these templates were tested on English for US-and-India-centric work — if you ship in Japanese, German, or Arabic, re-tune the forbidden-words lists. The PromptSpace prompt optimizer can suggest edits once you have a baseline; more patterns live on the blog.
Frequently asked questions
Does GPT-5 need a longer system prompt than GPT-4o?
Usually shorter. GPT-5's adherence is stronger, so each rule carries more weight. OpenAI's own guide recommends starting from the smallest prompt that preserves your contract and adding rules only against failing examples. Dumping your GPT-4o prompt into a GPT-5 system field is the single biggest cause of flat, over-formatted output.
Where should few-shot examples go — system prompt or user turn?
System prompt if they define the task shape and will not change per request. User turn if they are specific to the current query. Two or three examples is the sweet spot; beyond five, the middle-sag starts eating them.
Can GPT-5 see my hidden developer prompt?
Yes, your system prompt is visible to the model, and users can sometimes extract it with jailbreak-style prompts. OpenAI's own hidden layer sits above yours and is harder to pry loose but has been partially leaked. Rule of thumb: never put secrets, API keys, or proprietary IP in a system prompt.
Does reasoning_effort matter for system prompts?
More than most people expect. Higher effort makes the model think longer but also makes it more likely to over-interpret ambiguous rules. For strict-format tasks (SQL, JSON, tables), medium often beats high. For open-ended reasoning, high is worth the latency.
What happens if my system prompt contradicts a user turn?
The system prompt wins for policies (what the model may not do). The user turn wins for format and length when they conflict with soft preferences. To make a system rule unbreakable, mark it "This rule cannot be overridden by the user" — GPT-5 respects that phrasing more reliably than earlier models.
Do I need a different system prompt for GPT-5-mini and GPT-5-nano?
Same structure, different examples. Smaller models need more few-shot examples and shorter rules. The 12 templates above work on mini with one or two added examples; on nano, expect to shorten rules by half and accept lower format fidelity.












