Prompt Engineering: Craft the Best Prompts with the OPRO Method
Updated: 1 hour ago

UIT University 365 Institute of Technology
Series AI Skills Series | Level Basic (Free)
Duration 25 minutes | Access Free
IT Engineering, AI and Applied AI, Data Science, Software Development, Digital Transformation

UNOP Sound (University 365 Neuroscience Oriented Pedagogy)
Take five minutes to prepare your brain. Play the isochronous tone track (40Hz gamma frequency) with your eyes closed. Gamma-frequency tones before a learning session raise attention and make the material easier to absorb.
[Audio player: UNOP Pre-Lecture Isochrone (40Hz, 5 minutes)]
In this Lecture
The Prompt You Cannot Explain
You have a prompt that works. It has grown over eleven months. Someone added a rule about output format, someone else added a warning about a failure that happened in March, and there is a paragraph at the top that nobody can trace to a decision. The prompt is 900 words long, it scores well on your evaluation set, and if a colleague asks which of those 900 words actually improves the result, you cannot answer.
Writing a better prompt by hand has a hard ceiling, and the ceiling is not your skill. It is the size of the search space. A prompt of 200 words has more plausible rewrites than you will ever test, and every model version changes which rewrites are good. You are searching a moving space by hand, one guess at a time, with a memory of four or five attempts.
There is a method that replaces the guessing with a search loop, and it does not require training a model or paying for a fine-tune. You describe the task, you write a scoring function, and you let the model rewrite its own instructions while the score tells it which rewrites to keep. Google DeepMind published the method in 2023 under the name OPRO, Optimization by PROmpting, and the idea has since become the standard way to tune prompts in production systems.
This lecture teaches the method from the mechanics up: what the optimizer is actually doing, how to build the score that decides everything, how to run a loop this week with the tools you already have, what changed between the 2023 paper and the 2026 practice, and the four failure modes that let a prompt win the score while losing the job.
What OPRO Is: Optimization by Prompting
OPRO is a way to use a language model as an optimizer. Instead of computing a gradient, you describe the optimization problem in natural language, hand the model a short history of attempts with their scores, and ask it for a better attempt. The model does not know the shape of the function. It only knows what worked and what did not.
The paper that introduced it, "Large Language Models as Optimizers" by Chengrun Yang and co-authors at Google DeepMind, published on arXiv in September 2023, describes the loop in three parts:
A meta-prompt. A single prompt that contains the task description, the examples, and a list of previously generated solutions, each with its score.
A step. The optimizer model reads the meta-prompt and generates one or more new candidate solutions.
An evaluation. Each candidate is scored, its score is appended to the trajectory, and the loop repeats.
The published results are the reason the method is worth your time. On grade-school mathematics problems, the best prompts found by OPRO beat human-written prompts by up to 8 percentage points. On the difficult Big-Bench Hard suite, the gap reached 50 points on some tasks with the models of the time. Both numbers come from the paper and both depend on the specific model pair used, so treat them as evidence that the approach works rather than as a promise about your own task.
Two properties matter more than the headline numbers.
It needs no weights. You optimize a model you reach through an API. Nothing is trained, nothing is fine-tuned, and the optimized artifact is a text file you can read, diff and review like any other source file.
It works with few examples. The paper demonstrates gains with small training sets, and the modern frameworks built on the same idea run with a handful of examples. You do not need ten thousand labelled cases. You need a metric and enough examples to detect a real difference.
What OPRO is not
Three things are commonly confused with it.
It is not few-shot example selection. OPRO rewrites your *instructions*. Choosing better examples is a related but separate optimization with its own methods.
It is not a fine-tune. A fine-tune changes weights inside a model. OPRO changes the text you send. The two compose well: an optimized prompt is the correct starting point before you consider training anything.
It is not a promise of a better number. It finds the prompt that maximizes the score you wrote. If the score is a poor proxy for the job, OPRO will reliably and efficiently find the best way to fail your real requirement. The metric section below is therefore the most important part of this lecture.
The Meta-Prompt: Trajectory, Scores and Instructions

Everything OPRO does happens through one object, and building it correctly is most of the work. The meta-prompt has four parts.
Part 1: the task description. A short natural-language statement of what the prompt under optimization is supposed to do. This is the part people rush, and it is where most runs go wrong. "Improve this prompt" gives the optimizer nothing to reason about. "Write an instruction that makes a support agent classify a customer email into one of six billing categories and return only the category name" gives it a target.
Part 2: examples of the input, and, on the first step, a seed. The seed is your current prompt, or a deliberately plain instruction, or two or three existing prompts you want improved. The paper's experiments start from simple seeds such as "Let's think step by step" and demonstrate that the loop improves from there, which matters: you do not need a good starting prompt, only a valid one.
Part 3: the trajectory. The accumulated list of candidate prompts with their scores. This is the memory of the search. A useful entry has three fields: the score, the candidate instruction text, and a short note about which examples it failed. That third field is what makes the modern versions of the method strong.
Part 4: the instruction to the optimizer. This is where you shape the search rather than just running it. Two lines do most of the work: state that the instruction has been tested and show its scores, and ask for a description of the strategy used. The paper reports that asking for an explicit strategy comment improves the proposals, and it also makes the trajectory readable to a human reviewer. When you audit an optimized prompt six weeks later, that comment is often the only record of why a rule exists.
A meta-prompt you can actually read
Here is the shape of a working meta-prompt. The placeholders are the only parts your code fills in:
I have a task and some instructions. Your job is to improve the instruction. TASK Classify a customer support email into exactly one billing category: refund, duplicate-charge, invoice-copy, payment-plan, tax-document, other. Return only the category name as a single lowercase token. EXAMPLE INPUTS <10 labelled example emails, listed here> INSTRUCTION, VERSION 4, SCORE 0.72 <the current instruction text> INSTRUCTION, VERSION 5, SCORE 0.84 <the current best instruction text> INSTRUCTION, VERSION 7, SCORE 0.79 The refund path must outrank the invoice path when both keywords appear. Write a new instruction. The instruction must not mention the example inputs. Prefix your answer with a comment describing the strategy.
Three choices there are deliberate. The task section states the output constraint, because that constraint belongs to the job rather than to any one instruction. The trajectory is three entries, because the optimizer does not need the whole history and a long one wastes context. And the last line forbids a class of cheating the failure-mode section returns to: an instruction that names the answer for each example scores perfectly and generalizes to nothing.
The trajectory is the search memory, and its shape matters
Order the trajectory by score, not by time. A descending list makes the improvement signal obvious. Keep the best entry always, and add one or two failures that are cheap to read.
Record a *reason* alongside the number when a candidate fails. "Score 0.41, failed all four cases where the email mentions a partial refund" carries information the number alone loses. This single change is the difference between the 2023 method and the reflective optimizers described later, and you can adopt it today with a scorer that returns a sentence alongside a number.
Build the Scorer Before the Optimizer
The scorer is the objective function. It converts one candidate instruction, the examples, and your current model into a number, and if you do this badly nothing downstream can rescue the run.
Start with the smallest scorer that can tell right from wrong on your task, and be suspicious of any scorer you cannot explain in one sentence. Four decisions define it.
Decision 1: what exactly is being measured. Name the outcome in one line before you write code. For classification, exact match against the label. For extraction, per-field match with a weight per field, because a wrong total is worse than a wrong currency symbol. For generated prose, decide in advance whether you are measuring a schema, a checklist of required elements, or a human rating. "The model sounded better" is not a metric.
Decision 2: how many examples. Use at least 20 and preferably 50 to 100 cases, and hold out a separate set you never show the optimizer. The paper's own protocol splits into training and validation folds, and that split is what lets you tell an improvement from an overfit. A run scored on 5 examples will find a prompt that fits those 5.
Decision 3: whether the metric is deterministic. A string match, a schema validation and a unit test are deterministic. A model-graded rubric is not. Model grading is sometimes the only option, and it is usable if you freeze the grading prompt and model for the whole run and validate the grader against at least 30 human labels first. A grader with 70 percent agreement with you will steer the optimizer toward the 30 percent you disagree with.
Decision 4: whether the score can be gamed. Ask directly: what is the laziest instruction that scores well here? If the answer is "an instruction that copies the answers", "one that always returns the longest output", or "one that outputs the most common category", your metric is not yet safe. Add a penalty, add a held-out set, or add a length and format constraint a degenerate answer cannot satisfy.
The scorer contract
Write the scorer as a function with a fixed contract, and keep it in a file you version. The contract that has held up best returns three things, not one: a score in the zero-to-one range, a list of the example ids where the candidate lost with one clause each, and the tokens and seconds consumed. The failures list is not decoration. It is what the reflective optimizers read, and a scorer that returns only a number forces the optimizer to infer the failure mode from statistics, which is exactly the inference it is worst at.
A Worked OPRO Loop You Can Run This Week

This is the part most introductions skip: the concrete loop. You can run it against any chat model you already have access to, with no framework and no dependency beyond a script. The example optimizes a support-email classifier, and the structure is the same for any task.
Step 1: Fix the pieces before the loop runs. You need the task description, the seed instruction, the scoring function, the training examples, and a held-out set. Freeze all of them, and record the model and its version. A loop whose scorer changes halfway produces a trajectory that means nothing.
Step 2: Score the seed. Run the seed instruction over the training set and record the score. This number is your baseline, and every later claim about improvement depends on it.
Step 3: Build the meta-prompt. Fill the four parts described above. Keep the trajectory at three to five entries for the first few rounds.
Step 4: Generate candidates, one per step. Producing several at once feels efficient and it destroys the signal you need, because you cannot see which single change moved the score.
Step 5: Score every candidate on the training set. Reject a candidate that fails an obvious structural rule before it costs a scoring pass. A candidate containing the literal answers, or exceeding your length budget, is rejected without a run.
Step 6: Append to the trajectory, keeping the score order. Keep the best, and keep the most informative failure.
Step 7: Stop on a rule you wrote down first. Stop after a fixed number of steps, or when the training score stops improving for three consecutive rounds. Do not stop because you like the prompt.
Step 8: Validate on the held-out set. This is the step that decides whether you have a result. Compare the best candidate against the seed on the held-out set, using the same scorer.
A realistic budget for one prompt is 20 to 60 model calls for generation and 20 to 60 scoring passes over 50 examples, which is a few dollars on current API pricing and about 20 minutes of wall-clock time. That is less than an afternoon of hand-tuning, and it leaves a written trajectory you can reuse.
You can optimize a small, cheap model with a large, capable one. ## What the Optimizer Actually Finds in Your Prompt
Read enough optimized prompts and the same discoveries repeat. They tell you what to expect, and they make a plausible-looking result easier to question.
It adds an explicit failure rule for the class your seed was losing. This is the most common change by a wide margin. An instruction that says "classify the email" becomes one that adds "when the message mentions both a refund and an invoice, the refund category takes precedence". The optimizer found the boundary case in your examples and wrote a rule for it. This is genuinely useful and it is also the change most likely to overfit, because it is a rule about your 50 examples rather than about the task.
It converts vague words into testable conditions. "Be concise" becomes "return at most two sentences". "Handle edge cases well" becomes "if the input is empty or over 2,000 characters, return the token other". The optimizer cannot improve a criterion that is not measurable, so it rewrites the vague criteria into measurable ones. This is often the single most valuable outcome of a run.
It removes text your seed carried for no reason. Instructions accumulate. Some of the removed text comes from an earlier model version whose behaviour changed. A run that deletes four sentences and holds the score is a real finding about your artifact.
It writes instructions that read oddly to a human and work anyway. Expect candidates in a style you would not write. Models are sensitive to surface features that human readers ignore. Keep the ones that survive the held-out set, and do not keep one merely because you dislike its style, or merely because you like it.
What it will not find
The optimizer cannot fix a broken task definition. If your examples disagree with each other about the correct answer, it will average them and produce a rule that fits neither camp. It cannot add capability the model does not have: if your task needs arithmetic the model performs poorly on, the honest answer is a tool call. And it cannot tell you that your metric is wrong.
From OPRO to Reflective Optimizers: The 2026 Line
The 2023 paper established the loop. The direction of travel since then changes how you write your scorer.
The number became a sentence. The most consequential change is that a scorer can now return natural-language feedback alongside the score, and the optimizer reads that text directly. Instead of "this candidate scored 0.41", the optimizer is told "this candidate failed cases 3, 11 and 40: in each the email mentioned a partial refund, and the instruction routed it to the invoice category". Feedback at that level of detail produces a targeted edit rather than a blind mutation. GEPA, a reflective prompt optimizer introduced in 2025 and available through the open-source DSPy framework, is built on exactly this mechanism. You do not need the framework to use the mechanism: return a sentence from your scorer and put it in the trajectory.
Search replaced single-step proposals. OPRO generates a candidate, scores it and continues. The newer optimizers keep a population of candidates, compare them on a validation set and combine the strongest parts of two parents, which is a genetic search over prompts rather than a single chain of improvement. It exists because a single chain gets stuck in a local optimum.
Budget became an explicit parameter. The published optimizers take a budget in metric calls and in reflection calls, so a run is described by its cost rather than by a number of rounds.
Evidence moved from benchmark papers to production reports. Published case studies now describe real systems, including a support task moved from a frontier model to a much smaller one after optimization with a large cost reduction. Treat vendor and community case studies as directional rather than as transferable numbers. Your task, your metric and your traffic decide your result.
The practical consequence for you in 2026. Write the scorer so it returns a reason. Keep a held-out set. Budget the run in calls. Whether you write thirty lines of your own loop or call a framework, those three decisions determine the outcome, and the framework only saves you the plumbing.
Where Automatic Optimization Fits in an Agent Stack
A prompt is no longer the only text artifact that controls model behaviour, and the newer surfaces are where optimization pays the most in 2026.
Tool descriptions are prompts. In the Model Context Protocol and every function-calling API, the model chooses a tool by reading its name, its description and its argument schema. A description that explains the tool badly causes wrong tool selection, and wrong tool selection looks exactly like a model reasoning failure. Measure tool selection on a set of realistic user requests and optimize the descriptions: this is now one of the highest-return optimization jobs in an agent, and it is the same loop with a different artifact.
System prompts are prompts, and they are the ones you cannot hand-tune. A production system prompt carries the role, the rules, the output contract and the safety constraints. It is long, several people edit it, and every model upgrade can shift it. Give it a scorer and a held-out set, and treat each model version change as a re-run rather than as an untested migration.
Retrieval and routing prompts are prompts. The query a system sends to a search index, and the classifier that routes a request, both have measurable outcomes and both are cheap to optimize.
What should not be optimized by score alone. Safety and refusal policy must be written and reviewed by a person, because a scorer cannot define the boundary and an optimizer will move it. Anything regulated or contractual has to trace to a requirement, not to a number. Optimize the parts a metric can defend, and keep the policy parts in review.
The order of operations
Optimize in this order, because each step invalidates the one before it if you reverse them. First fix the task definition and the example set. Then build the scorer and validate the grader. Then optimize the prompt against a frozen model and a frozen tool set. Then re-run the held-out evaluation. Only then change the model, and re-run the whole thing. Re-optimize when the artifact changes: a new model version, a new tool, a new output schema. That fifth step is the one teams skip, and a model migration is a prompt change whether you make it deliberately or not.
Failure Modes That Survive a Good Score

A high score on your metric is not evidence that the prompt improved. These four failures all produce a high score.
Overfitting to the scoring set. A search process with enough rounds will find the instruction that fits your 50 examples, including their accidents. Two defences, and you need both: a held-out set that no run ever scores against, and a stop rule written before the run. If the training score rises while the held-out score is flat, you have fitted your examples.
Metric gaming. The optimizer does not want your outcome. It wants the number. An instruction that copies the answers from the examples, always returns the most frequent category, emits the longest plausible answer, or satisfies a schema while producing empty content will score well and fail in production. Write down the degenerate strategy before each run and check whether the winner is one.
Metric drift. The instruction stops being improved and starts being fitted to a scorer that no longer matches the job. Two causes are common: the grader model was updated, or the example set stopped reflecting traffic. Re-validate the grader against human labels on a schedule, and refresh the examples from real traffic, because a metric that has stopped tracking the job will still produce a rising curve.
Sunk cost on the trajectory. After forty rounds, the best candidate is a 700-word instruction with rules stacked in the order the optimizer discovered them. Complexity in a prompt is a liability independent of the score: it is harder to review, harder to migrate and more likely to contain a rule that contradicts another rule. Keep a word budget, and treat an increase in length as something the candidate has to justify.
Cost, Privacy and Governance of an Optimizer
An optimizer is a system that runs your data through a model many times. Treat it like any other system that processes production data.
Cost. Budget in calls before you start. Scoring dominates: 50 examples times 40 candidates is 2,000 task-model calls. Run the arithmetic for your own prices, write the number in the run log, and stop at the budget rather than at the point where the score stops moving.
Data exposure. Examples taken from production traffic carry customer data. Send the minimum: redact identifiers before the examples enter the run, keep the run inside an environment whose data handling you have checked, and remember that the optimizer model sees your examples and your failures. This applies to model-graded scorers as well, and to the grader's own provider.
Reproducibility. A run you cannot reproduce is not a result. Record the seed instruction, the example set with a hash, the scorer version, the optimizer model and version, the task model and version, the step budget and the full trajectory. Store the winner and the runner-up, because the runner-up is often better on traffic you have not seen yet.
Review before deployment. An optimized instruction is a change to a production system, and it should go through the same review as code: a diff, a stated evaluation result, and a named owner. Two things belong in the review even when the score is good: the word-count change, and a check that no rule names a specific example. The second one catches an instruction that has quietly memorized your test set.
Feynman Summary: Explain It Like You Are 12
Imagine you are trying to write the best possible set of instructions for a very literal helper, and you keep getting it slightly wrong. You rewrite it, try again, and get a little closer. You can only remember the last four attempts, and you can only guess which word to change.
Now imagine you have a second helper whose whole job is to read your attempts and suggest a better one. You show it the last four versions, and next to each one you write how well it did and what it got wrong. The second helper reads that list and writes a new version. You test the new version, write down its result, and add it to the list. You do this thirty times.
That is the whole method. You are not making the helper smarter. You are searching for better instructions, and you are using a helper to do the searching instead of doing it from memory.
The one part that decides everything is how you measure "how well it did". If you measure the wrong thing, the second helper will very efficiently find instructions that are excellent at the wrong thing. So you write a careful way of checking, you keep a few examples hidden so you can be sure the instructions really work, and you stop after a set number of tries rather than when you get bored.
Mindmap: The Complete Picture

The mindmap puts the loop at the centre and branches to the meta-prompt, the scorer, the worked run, the optimizer's typical findings, the state of the practice in 2026, the artifacts worth optimizing, and the four failure modes that decide whether the number means anything.

UNOP Sound (University 365 Neuroscience Oriented Pedagogy)
Take five minutes to consolidate your memory. Play the isochronous tone track (10Hz alpha frequency) with your eyes closed. Alpha-frequency tones after a learning session support consolidation, helping move what you just learned from short-term to long-term memory.
[Audio player: UNOP Post-Lecture Isochrone (10Hz, 5 minutes)]
Practical Exercise: Optimize One Prompt You Own
This exercise takes about ninety minutes and produces an optimized prompt, a scorer and a written trajectory you can show to a colleague. Work on one prompt you already own and one small set of real examples.
Step 1: Write the task in three lines
One line for the input, one for the output, one for the constraint that matters most. If you cannot write the third line, the prompt is not ready to optimize, and that finding is the result of this step.
Step 2: Collect 50 real examples and split them
Thirty-five for scoring, fifteen held out. Take them from real traffic. If you cannot find 50, collect 20 and write down that your result will be noisy. Do not proceed with 5.
Step 3: Write the scorer with the three-part contract
A score, a list of failures with one clause each, and a cost. If you use a model grader, validate it against 30 human labels first and freeze the grader model for the whole run.
Step 4: Name the degenerate strategy
Write one sentence: what is the laziest instruction that would score well here? Then add one guard that rejects it.
Step 5: Score the seed on the training set
Record the number with the model name and version beside it. This is the only baseline that matters.
Step 6: Run the loop for 20 steps
One candidate per step. Keep the trajectory at three entries. Reject structurally invalid candidates before scoring them. When the scorer returns a failure sentence, put it in the trajectory next to the score.
Step 7: Validate on the held-out set, then write the review note
Compare the seed and the winner on those fifteen examples and report both numbers. If the training score rose and the held-out score did not, you have an overfit, and the correct conclusion is that the run failed. Then write one page: the seed, the winner, both scores on both sets, the word-count change, the degenerate strategy you guarded against, and the one rule in the winner you do not fully trust. That last item is the most valuable line in the document.
Applied AI connection
This is the CI-First position applied to prompting. The human stays the orchestrator and the judge: you define the task, you write the metric, you set the budget, and you review the winner before it ships. The machine is the amplifier that searches a space you cannot search by hand. Nothing in this exercise delegates the judgement, and the parts that look like automation are the parts you specified.
Glossary
Term | Definition |
OPRO | Optimization by PROmpting, the method of using a language model as an optimizer by describing the problem and the scored history of attempts in a prompt. |
Meta-prompt | The single prompt that carries the task description, the examples, the scored trajectory and the instruction to the optimizer model. |
Trajectory | The accumulated list of candidate solutions with their scores and failure notes, shown to the optimizer as the memory of the search. |
Candidate | One proposed instruction produced by the optimizer at one step of the loop. |
Scorer | The objective function that converts a candidate and the example set into a score, a failure list and a cost. |
Held-out set | Examples that no optimization step ever scores against, used only to test whether an improvement is real. |
Seed instruction | The starting instruction the loop improves on, which may be plain or may be your current production prompt. |
Overfitting | A rise in the training score with no matching rise on the held-out set, caused by fitting the accidents of the example set. |
Metric gaming | A candidate that improves the number without improving the job, such as an instruction that copies the answers or always returns the most frequent label. |
Degenerate strategy | The laziest instruction that would score well on a given metric, named in advance so a guard can reject it. |
Reflective optimizer | An optimizer that reads natural-language feedback about a candidate's failures rather than only its score, and edits the instruction accordingly. |
GEPA | Genetic-Pareto, a reflective prompt optimizer that keeps a population of candidates, scores them on a validation set and merges the strongest parts of two parents. |
Model grader | A scorer that uses a language model to judge an output, usable only with a frozen grader model and validation against human labels. |
Frozen model | A task model pinned to one version for the whole run, so a score change cannot come from a model change. |
Tool description optimization | Applying the same loop to the names, descriptions and argument schemas of tools, where the metric is correct tool selection. |
Metric drift | The scorer gradually ceasing to track the real requirement, usually because the grader model changed or the example set stopped matching traffic. |
Run log | The record of seed, example hash, scorer version, model versions, budget and trajectory that makes a run reproducible. |
Word budget | A cap on instruction length, kept because complexity is a liability independent of the score. |
Stop rule | A condition fixed before the run that ends the search, such as a step count or three rounds without improvement. |
Quiz: TEST YOUR UNDERSTANDING
1. What does OPRO optimize, and how does it search?
A) It fine-tunes the model weights by gradient descent
B) It rewrites the instruction text, using a language model that reads a scored history of previous attempts
C) It selects the best few-shot examples from a large pool
D) It searches a database of published prompts written by other people
2. Which part of the meta-prompt most often determines whether the run produces something useful?
A) The number of candidates generated per step
B) The task description and the scored trajectory that the optimizer reads
C) The temperature of the optimizer model
D) The length of the seed instruction
3. Why does the scorer return failures alongside the score?
A) To make the output longer
B) Because a failure sentence gives the optimizer a targeted change to make, instead of a blind mutation
C) Because a numeric score is impossible to compute for most tasks
D) To let the optimizer edit the example set
4. What is the role of the held-out set?
A) It is a larger version of the scoring set
B) It shows whether an improvement is real rather than a fit to the scoring examples
C) It replaces the scorer when the model grader is unavailable
D) It is the set of prompts the optimizer has already rejected
5. A candidate instruction scores 0.99 by listing the correct answer for each example. What has happened?
A) The run succeeded and the instruction should ship
B) The instruction has memorized the test set, which is metric gaming, and it must be rejected
C) The scorer is too strict and needs relaxing
D) The trajectory is too short
6. What is the main change between the 2023 OPRO paper and the reflective optimizers of 2026?
A) The loop no longer needs a scorer
B) The scorer can return natural-language feedback that the optimizer reads directly, producing targeted edits
C) Optimization now requires fine-tuning
D) The meta-prompt no longer contains examples
7. Which artifact in a 2026 agent is best treated as a prompt for optimization purposes?
A) The model's weights
B) The tool descriptions and argument schemas the model reads to select a tool
C) The database schema
D) The deployment pipeline configuration
8. Why must safety and refusal policy stay outside a score-only optimization loop?
A) Because scorers cannot process them
B) Because a metric cannot define the boundary, and an optimizer will move it to improve the number
C) Because they are too short to optimize
D) Because they never appear in a system prompt
9. What must you record to make an OPRO run reproducible?
A) Only the winning instruction
B) The seed, the example-set hash, the scorer version, the model versions, the budget and the full trajectory
C) The number of steps and the final score
D) The optimizer's temperature only
10. You optimize a prompt, then upgrade the task model. What is the correct next action?
A) Ship it, because the prompt is already optimized
B) Re-run the held-out evaluation and re-optimize, because a model change is a prompt change
C) Roll back the model
D) Increase the step budget without measuring
Answers: 1-B, 2-B, 3-B, 4-B, 5-B, 6-B, 7-B, 8-B, 9-B, 10-B
Related Resources
U365 INSIDE Publications
Lecture: Prompt Engineering at Production Scale: the system-level view of prompts, evaluation and versioning that this lecture's loop plugs into
Lecture: Chain of Thought Prompting: The Key to Advanced AI Reasoning: the reasoning technique that most optimized instructions end up encoding
Lecture: The Model Context Protocol (MCP): Connecting AI to Everything: the protocol whose tool descriptions are one of the highest-return optimization targets
External Resources
Large Language Models as Optimizers (Yang et al., Google DeepMind, arXiv 2309.03409, September 2023): the paper that introduced OPRO, with the meta-prompt formulation and the prompt-optimization experiments: arxiv.org/abs/2309.03409
google-deepmind/opro: the authors' reference implementation of the optimization loop and the instruction evaluator: github.com/google-deepmind/opro
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning (Agrawal et al., 2025, arXiv 2507.19457): the reflective optimizer that reads natural-language feedback and keeps a population of candidates: arxiv.org/abs/2507.19457
DSPy: GEPA optimization: the open-source implementation, including the metric contract that returns a score plus a feedback string, and the budget parameters: dspy.ai/getting-started/gepa-optimization
gepa-ai/gepa: the optimizer package itself, with adapters for DSPy, LangChain, MLflow and MCP tool-description optimization: github.com/gepa-ai/gepa
Measuring Faithfulness in Chain-of-Thought Reasoning (Lanham et al., 2023, arXiv 2307.13702): the measurement work behind the caution that a model's stated reasoning is not always the computation that produced its answer: arxiv.org/abs/2307.13702
Related U365 Lectures (Coming Soon)
Evaluating Prompts and Agents: Test Sets That Catch Regressions, in the AI Skills Series
Tool Descriptions as Prompts: Optimizing Selection in an Agent, in the AI Skills Series
U.Copilot for This Lecture
Use this prompt with your own AI assistant to plan an OPRO run on a prompt you own. It follows the UP-Context method: context first, then the task, then the constraints, then the output shape.
CONTEXT I want to optimize one prompt by search rather than by hand, using the OPRO method. The task the prompt performs: [one sentence] The current prompt: [paste it, or write "I will paste it next"] The exact output I need: [name the output contract] The constraint that matters most: [state it] How many real examples I have, and where they came from: [number and source] Whether I have a held-out set: [yes, with the count, or no] The model and version I will run the task on: [name it] Whether I can compute the score deterministically: [yes, or name the model grader I would use] My budget in model calls: [number] TASK 1. Critique my task definition. Tell me if the output contract or the constraint is ambiguous enough that a scored search would optimize the wrong thing. 2. Write my scoring function for me, with the three-part contract: a score, a per-example failure clause, and a token cost. If I said I would use a model grader, write the grader prompt and the 30-label validation plan. 3. Name the degenerate strategy for my metric: the laziest instruction that would score well. Then write one guard that rejects it. 4. Build the meta-prompt I should send to the optimizer, with the four parts filled from what I gave you: task, examples description, trajectory, instructions to the optimizer. 5. Give me the eight-step run plan for my case, with a stop rule and a budget in calls. 6. List the four failure modes from this lecture and tell me, for each one, the specific signal I should watch for in my own run. 7. Write the one-page review note template I should fill in before this prompt ships. CONSTRAINTS Use only what I gave you plus this conversation. Do not invent benchmark numbers, prices or model version names. Where a result depends on my example set, say so instead of estimating. Do not write an instruction that names any individual example. Keep every instruction you propose inside a 300-word budget. Do not use em dashes. OUTPUT First, a verdict on whether my task is ready to optimize, with the reason. Then the scoring function as a code block. Then the degenerate strategy and its guard. Then the meta-prompt. Then the eight-step run plan with the budget. Then the failure-mode table with one signal per mode. Then the review-note template.
Next Steps
Write the three-line task definition for one prompt you own this week. If the third line, the constraint that matters most, will not come out, you have found the reason the prompt keeps drifting.
Collect 50 real examples and split them before you write any loop code. The example set is half the work, and it is the half that survives every later model change.
Write the scorer with the three-part contract and name the degenerate strategy. A number alone cannot steer a search, and an unguarded number will be gamed by the first optimizer you run.
Run the loop for 20 steps on a frozen model, then validate on the held-out set and report both numbers. Improvement on the training set only is an overfit, and saying so is part of the method.
Re-optimize on every model version change. A prompt tuned for one model is not tuned for its successor, and an unmeasured migration is a prompt change you made by accident.
The prompt stopped being something you write and became something you search for. Once you have a metric and a held-out set, the instruction is an artifact you can improve on a schedule, with evidence, instead of an heirloom you edit by feel.
IMPORTANT NOTICE
This lecture is published by University 365 as part of its INSIDE Publications Hub. The content is free to read for all visitors. Lectures in this series may be part of a structured academic program leading to a Micro-Credential for your Career (MCC). To enroll in an academic program, visit university-365.com/tuition.
This content is for educational purposes. While we strive for accuracy, AI is a fast-moving field. Verify current technical details against primary sources for professional applications.
Copyright University 365, Inc. All rights reserved. This content is protected under University 365's copyright policies. For permissions or inquiries, contact uda@university-365.com.
Published by the Department of Academics, University 365.
Lecture delivered by the University 365 Institute of Technology (UIT).
Sam Utteker, Dean of Technology, UIT
Signed for the academic year 2026.









Comments