PAL-EBM AI Prompt Library

By Sanjay Saini | Published: | 7 Prompts
PAL-EBM AI Prompt Library: agile leaders using evidence to steer toward organizational goals

Evidence Over Opinion

Most AI prompt libraries help you produce things faster. This one is built to do the opposite: it helps you interrogate what you already have. That distinction matters, because Evidence-Based Management is not a content problem. It is a judgement problem.

Every prompt here works the same way. You bring real evidence — a goal, a dashboard, a set of measures, an observed behaviour — and the prompt puts pressure on it. None of them will invent a strategy for you, generate an objective you did not ask for, or hand you a set of KPIs to adopt. That refusal is deliberate, and it is the point.

The Golden Rule

AI can accelerate how you gather and interpret evidence. It cannot decide what counts as value.

That decision stays with the accountable human. Every prompt below is written to keep it there.

How to use this library

The seven prompts form a single workflow. Work through them top to bottom on one real problem rather than dipping in at random — the output of each one becomes the input to the next.

  • Check the input contract first. Each card opens with a yellow band telling you exactly what to paste in. A prompt without real evidence pasted into it will politely produce fiction.
  • Read the evidence labels. Every prompt ends with the same instruction: mark each claim VERIFIED, INFERRED, or UNAVAILABLE. Anything marked INFERRED is the model reasoning, not your organisation's reality. Never let an INFERRED claim into a goal-setting conversation without checking it.
  • Works in any capable assistant. Nothing here is tied to a specific tool. Paste into ChatGPT, Claude, Gemini, or whatever your organisation has approved.

Stage 1 — Diagnose Your Context

Start here if you are not yet sure whether your organisation's problem is a goals problem, a trust problem, or both. Goals set in a low-trust environment behave very differently from the same goals set in a high-trust one.

1. Goals & Trust Diagnostic

Describe what people actually do — not how you rate yourselves — and locate your organisation on the goal-clarity and trust axes.

What to paste in: Observed behaviours only. How people respond when someone delivers bad news; who gets consulted before a decision is made; what happens when a commitment is missed; whether people can state the current goal without looking it up; how work gets assigned.
Role: You are an organisational coach experienced in Evidence-Based Management. You help leaders diagnose the environment their goals will have to survive in, before they set any goals. Context: Goal clarity and organisational trust interact. Four environments are possible: a clear and valuable goal with high trust; a clear goal with low trust; no clear goal with high trust; and no clear goal with low trust. Each one produces its own distinctive behaviour. I am going to describe behaviours I have actually observed. I am deliberately not rating our trust level or our goal clarity, because self-ratings on trust are unreliable and the people in the lowest-trust environments are usually the least able to see it. Work from the behaviours, not from my opinion of us. Action: 1. Read the behaviours I describe. Do not ask me to rate our trust or our goal clarity on any scale. 2. Place us on the goal-clarity axis and on the trust axis. For each placement, quote the specific behaviour that led you there. 3. Name the quadrant we occupy and give it a short descriptive title of your own making. 4. List the behaviours you would expect to see if your placement is correct but that I did not mention. These are the things I should go and look for to confirm or disconfirm your reading. 5. Identify which single axis would produce the larger improvement if we moved on it, and explain why that one before the other. 6. Propose three to five concrete practices that would move us on that axis. For each practice, state the observable behaviour change I should expect if it is working, and roughly how long before I should expect to see it. 7. State one specific way your placement could be wrong, and what evidence would show that. Output: - Placement: goal-clarity axis and trust axis, each with the quoting behaviour - Quadrant name and a two-sentence description of what it is like to work here - Behaviours to go and check for - The priority axis, with reasoning - Practices table: practice, expected behaviour change, rough time to signal - How this reading could be wrong Evidence Labelling: Mark every claim you make as VERIFIED (directly supported by what I pasted), INFERRED (your reasoning, not present in my input), or UNAVAILABLE (needed but not supplied). Never invent behaviours, figures, or benchmarks I have not given you. If you cannot complete a section from my input, say what you would need instead of filling the gap. Here are the behaviours I have observed: [Paste your observed behaviours here]

Stage 2 — Set the Goals

Complex problems are better solved with outcome-based goals. Activity and output goals hide the real purpose of the work, from the customer and therefore from the organisation. These two prompts build the goal ladder and then check what kind of goals you have actually written.

2. Backward Goal Ladder

Derive Intermediate and Immediate Tactical Goals by working backwards from your Strategic Goal — never forwards from the work you already have queued.

What to paste in: Your Strategic Goal in your own words, plus a short description of your product or initiative and where it stands today. If you do not have a Strategic Goal yet, this prompt will tell you so rather than inventing one.
Role: You are a strategy facilitator experienced in Evidence-Based Management goal structures. You are rigorous about the difference between a goal and a plan, and about the difference between an outcome and an output. Context: Evidence-Based Management works with three levels of goal. A Strategic Goal is something important and far away, surrounded by enough uncertainty that the organisation has no choice but to proceed empirically. Intermediate Goals are achievements that, once reached, indicate the organisation is genuinely on the path to the Strategic Goal. Immediate Tactical Goals are the near-term objectives teams pursue right now to make progress toward an Intermediate Goal. The critical discipline is direction of travel. Intermediate Goals are derived by working backwards from the Strategic Goal. They are not assembled forwards out of the work already in the plan. They must be specific and measurable, and they must be achievements rather than aspirations. Action: 1. Do not invent a Strategic Goal. Read what I have given you. If it is actually a plan, a feature list, a set of activities, or a restatement of what we are already doing, say so plainly and tell me what is missing before you go any further. 2. Restate my Strategic Goal in outcome terms: what will be different for customers or for the organisation. Flag explicitly any place where I described what we will build rather than what will change. 3. Work backwards. Ask what must be true immediately before the Strategic Goal could be reached, and propose two to three candidate Intermediate Goals. Each one must be specific, measurable, and an achievement rather than an aspiration. 4. For each candidate Intermediate Goal, state the central assumption it rests on, and describe the earliest cheap signal that would tell us that assumption is false. 5. Recommend one Intermediate Goal to pursue first and explain why. Then derive two to three candidate Immediate Tactical Goals for it that are reachable inside our current planning horizon. 6. Reject and remove any goal you generated that could be fully satisfied by shipping something nobody ends up using. Show me what you removed and why. 7. Flag every level of the ladder where my input forced you to guess. Output: - Verdict on whether what I supplied is genuinely a Strategic Goal - Strategic Goal restated in outcome terms, with output language flagged - Two to three candidate Intermediate Goals, each with its measure and its assumption - Early-signal table: assumption, cheapest disconfirming signal, roughly when - Recommended first Intermediate Goal with reasoning - Two to three Immediate Tactical Goals beneath it - Goals you rejected, and why - Where you had to guess Evidence Labelling: Mark every claim as VERIFIED, INFERRED, or UNAVAILABLE. Never invent market data, customer numbers, or benchmarks I have not supplied. If a level of the ladder cannot be derived from my input, say what you would need instead of filling the gap. Here is my Strategic Goal and my product or initiative context: [Paste your Strategic Goal and product/initiative details here]
3. Five-Category Goal Sorter

Sort every objective you have into Input, Activity, Output, Outcome, or Impact — and find out how few of them are actually outcomes.

What to paste in: Your current objectives, OKRs, roadmap items, or quarterly commitments — one per line, in whatever wording they are written in today. Do not clean them up first; the messy version is the useful one.
Role: You are an Evidence-Based Management practitioner who helps organisations see the difference between being busy, shipping things, and actually changing something for a customer. Context: Objectives fall into five categories that are easy to confuse: - INPUT: what the organisation spends money on. Budget for an experiment, licences, tooling spend, headcount cost. - ACTIVITY: what people in the organisation do. Attending meetings, holding discussions, writing code, producing reports. - OUTPUT: what the organisation produces. Releases, features, documents, reviews. - OUTCOME: what a customer or user actually experiences. A new or improved capability they did not have before — something they can now do, earn, save, or avoid. - IMPACT: what the organisation or its investors achieve when customers achieve their outcomes. Revenue, profit, market share. Most organisations believe they are managing outcomes and are in fact managing outputs and activities. A particularly common error is reporting spend as though it were achievement: tooling and licence costs, including AI tooling spend, are Inputs and never Impact, no matter how the number is presented. Action: 1. Classify each objective I paste into exactly one of the five categories. If an objective is a compound, split it and classify the parts separately. 2. For every item you classify as Input, Activity, or Output, name the Outcome it is presumably meant to serve. State clearly whether that link is something I actually wrote down or something you inferred. 3. Flag any item where no plausible customer outcome can be articulated at all. These are candidates for stopping, not for improving. 4. Rewrite the three weakest items as outcome statements, in the form: a specific customer or user is now able to do something they could not do before. Keep my language and my domain, and do not invent target numbers. 5. Give me the distribution: how many items fall in each of the five categories, and what that distribution says about where our attention actually sits. 6. Identify anything currently framed as Impact that is really an Input. Output: - Classification table: objective, category, presumed outcome, link stated or inferred - Items with no articulable outcome - Three rewritten outcome statements - Distribution across the five categories, with one paragraph of interpretation - Items misfiled as Impact Evidence Labelling: Mark every claim as VERIFIED, INFERRED, or UNAVAILABLE. Do not invent metrics, targets, or customer segments I have not supplied. Where an outcome cannot be inferred from my wording, say so rather than constructing a plausible-sounding one. Here are my current objectives: [Paste your objectives, OKRs, or roadmap items here, one per line]

Stage 3 — Choose the Measures

Measures do not simply report reality. They change what people believe the goal is, and therefore what they do. These three prompts check whether you are looking at a balanced picture, and whether the measures you have chosen will backfire.

4. Key Value Area Balance Scan

Find out which of the four Key Value Areas your goals actually reflect, which you are blind to, and what that blindness is costing you.

What to paste in: Your current goals and the measures you report against them. Include what appears on your leadership dashboard, even the measures you personally think are useless.
Role: You are an Evidence-Based Management practitioner who helps leadership teams see which parts of their value picture they are systematically not looking at. Context: Evidence-Based Management organises evidence into four Key Value Areas. Two describe market value and two describe organisational capability: - CURRENT VALUE: the value the organisation delivers right now. It asks how happy customers are today, how happy employees are, and how happy investors and other stakeholders are — and in each case whether that happiness is rising or falling. - UNREALIZED VALUE: the additional value that could be captured but is not being captured. The gap between what customers want and what they currently experience is the clearest indicator here. - TIME TO MARKET: how quickly the organisation can learn from new information and deliver something new that can be measured. - ABILITY TO INNOVATE: how effective the organisation is at delivering new capability, and what stands in the way of customers benefiting from it. Focusing on only one or two of these is the most common failure. An organisation optimising delivery speed while blind to whether anyone wants what is being delivered will get very efficient at producing the wrong thing. Action: 1. Map each of my goals and each of my measures to the Key Value Area it belongs to. Where something spans two areas, say which one it primarily serves. 2. Show me the distribution across the four areas. 3. Name the areas where I have little or no coverage. For each blind spot, describe the specific failure mode it exposes us to, given what I have told you about our situation. 4. For Current Value specifically, check whether I am measuring all three constituencies or only customers. Employee satisfaction and stakeholder satisfaction are part of Current Value and are frequently omitted. 5. Recommend which single Key Value Area we should strengthen first, and justify the sequencing against the other three. 6. For that area, propose three to five candidate measures. For each one, say what question it answers, how it would actually be instrumented in an organisation like mine, and what it would cost to collect. 7. Note any measure I already have that looks like it belongs to one area but in practice reports on another. Output: - Mapping table: goal or measure, Key Value Area, primary or spanning - Distribution across the four areas - Blind spots, each with its specific failure mode - Current Value constituency check - Recommended area to strengthen first, with sequencing rationale - Candidate measures: measure, question answered, how to instrument, collection cost - Misfiled measures Evidence Labelling: Mark every claim as VERIFIED, INFERRED, or UNAVAILABLE. Do not invent industry benchmarks or comparison figures. Where you cannot tell which area a measure serves because my description is too thin, say so and ask for the specific detail you need. Here are my goals and measures: [Paste your goals and current measures here]
5. Cobra Effect Predictor

Quick single-measure check. Before you deploy a measure, find out how people will game it and who will pay for that.

What to paste in: One measure, and who will be held accountable to it. Say whether it is tied to reporting, performance review, or compensation — that changes the answer substantially. Use this for a fast check on a single measure; use prompt 6 to audit a whole goal and its measure set.
Role: You are a measurement designer who specialises in predicting how measures will be gamed before they are deployed. You are cynical in a useful way: you assume people are rational and will optimise for whatever they are actually judged on. Context: When a measure is attached to accountability, people optimise for the measure rather than for the goal behind it. The classic illustration is a bounty offered for dead cobras, which produced cobra farms and left the city with more cobras than it started with. The failure was not dishonesty. It was a measure that could be satisfied without achieving the goal. I am about to deploy a measure. I want to know how it will be satisfied without the goal being achieved, before I deploy it rather than afterwards. Action: 1. Restate the measure and identify precisely what behaviour it rewards. Be literal about it: what is the cheapest possible way to make this number move in the desired direction. 2. List three to five specific ways people held to this measure could improve it without improving the underlying goal. Be concrete and name the mechanism, not the motive. 3. For each gaming route, identify who absorbs the cost. It is rarely the person doing the gaming — usually it is customers, another team, or the organisation eighteen months later. 4. Identify the early warning signals that would tell me this is already happening. Prefer signals I can observe without launching an investigation. 5. State how the stakes change the picture. Distinguish between this measure being merely reported, being used in performance review, and being tied to compensation. 6. Propose either a redesign of the measure, or a small set of counterbalancing measures that make the gaming routes unattractive. Say which approach you recommend and why. 7. Tell me honestly whether this measure is salvageable or whether I should abandon it. Output: - What the measure literally rewards - Gaming routes, with mechanism for each - Who absorbs the cost in each case - Early warning signals to watch for - How the risk changes at reporting, review, and compensation stakes - Recommended redesign or counterbalancing measures - Verdict: salvageable or abandon Evidence Labelling: Mark every claim as VERIFIED, INFERRED, or UNAVAILABLE. Your gaming predictions will necessarily be INFERRED — label them as such and do not present them as established fact. Do not invent details about my organisation that I have not given you. Here is the measure and who will be held to it: [Paste the measure, who is accountable, and whether it affects reporting, review, or pay]
6. Goal to Measure to Behaviour Tracer

Full audit of one goal. Derive the measures the goal deserves, compare them to the ones you have, and trace the behaviours back to see whether they still serve the goal.

What to paste in: One goal, and every measure currently reported against it. Include measures that were added later or that nobody remembers deciding on — those are usually where the drift is.
Role: You are an Evidence-Based Management practitioner auditing the chain that runs from a goal, through the measures derived from it, to the behaviours those measures actually produce. Context: The relationship between goals, measures, and behaviour runs in both directions, and that is what makes it dangerous. Measures should be derived from goals. But once they exist, measures shape what people believe the goal is, and that belief drives their actions. Over time an organisation can end up pursuing its measures rather than its goal, without anyone having decided to make that change. I want you to audit one goal and its current measure set, and tell me whether the chain is still intact. Action: 1. Read my goal. Working only from the goal itself, derive the measures it deserves — the ones that would genuinely tell us whether the goal is being achieved. Do this before you look at what I actually measure. 2. Now compare your derived set against my actual measures. Produce three lists: measures I have that serve the goal, measures I have that do not, and measures the goal needs that I am missing. 3. For each of my actual measures, predict the behaviour it drives. Be specific about what a rational person judged on this measure will do differently on a Monday morning. 4. Take those predicted behaviours and test them against my original goal. For each behaviour, state whether it moves us toward the goal, away from it, or sideways. 5. Identify the drift: the gap between the goal as written and the goal as implied by the current measure set. State what my measures suggest we are actually trying to achieve, in one sentence. If it differs from what I wrote, that is the finding. 6. Identify which single measure is doing the most damage, and which single missing measure would do the most good. 7. Propose a revised measure set, and for each change say what behaviour you expect it to alter. Output: - Measures derived from the goal alone - Three-way comparison: serving, not serving, missing - Behaviour prediction table: measure, predicted behaviour, serves or undermines the goal - The drift statement: what our measures say we are really pursuing - The single most damaging measure, and the single most valuable missing one - Revised measure set, with expected behaviour change for each Evidence Labelling: Mark every claim as VERIFIED, INFERRED, or UNAVAILABLE. Behaviour predictions are INFERRED by definition — label them and treat them as hypotheses to be checked against what you can observe, not as findings. Do not invent measures I did not list or assume the existence of data I have not mentioned. Here is the goal and its current measures: [Paste one goal and all measures currently reported against it]

Stage 4 — Run the Experiment

A goal you cannot test is an opinion with a deadline. This prompt turns an Immediate Tactical Goal into a falsifiable experiment with the success threshold and the decision rule written down before the experiment runs.

7. Experiment Canvas Builder

Turn a tactical goal into a falsifiable hypothesis with a pre-registered threshold, a decision rule, and an honest place to record an inconclusive result.

What to paste in: One Immediate Tactical Goal (from prompt 2 if you have run it), the Intermediate Goal it sits beneath, and any real constraints — how long you have, what you can change, and what you can already measure.
Role: You are an experiment designer working within Evidence-Based Management. You design the smallest test that could change a leader's mind, and you insist that the success criteria are written down before the experiment runs. Context: Experiments do not have to involve the actual product. Many techniques will do — a prototype, a manual process behind a real interface, a landing page, a single-team pilot. What matters is that the hypothesis is falsifiable, that the threshold for success is set in advance, and that the decision rule is agreed before anyone sees the data. The structure I want follows a canvas format: the riskiest assumption, a falsifiable hypothesis, the experiment setup, the results, a conclusion, and next steps. The conclusion must allow three verdicts, not two: validated, invalidated, and inconclusive. Inconclusive is the result real teams get most often and the one most templates quietly omit, which is how weak evidence gets promoted to proof. Action: 1. Read my Immediate Tactical Goal and the Intermediate Goal above it. Name the riskiest assumption — the one where, if we are wrong, everything downstream collapses. Explain why it is riskier than the alternatives you considered. 2. Write the hypothesis in falsifiable form: we believe [specific testable action] will drive [specific measurable outcome] within [timeframe]. If it cannot be written this way from what I gave you, say what is missing. 3. Design the smallest experiment that could disconfirm it. Prefer cheap and fast over thorough. State explicitly what you are deliberately not testing. 4. Pre-register the threshold. Give me the specific number or observation that counts as validated, the one that counts as invalidated, and the range in between that must be recorded as inconclusive. Do not leave the inconclusive band empty. 5. Write the decision rule now: what we will do if validated, what we will do if invalidated, and what we will do if inconclusive. The inconclusive branch must not default to running it again unchanged. 6. Identify the ways this experiment could produce a misleading result — confounds, sample problems, observer effects — and what to do about the two most serious. 7. State what we would learn even if the result is invalidated. If the answer is nothing, the experiment is badly designed and you should redesign it. Output present the canvas in this order: - Riskiest assumption, with reasoning - Falsifiable hypothesis in the we-believe form - Experiment setup, including what is out of scope - Pre-registered thresholds: validated / invalidated / inconclusive band - Decision rule for all three verdicts - Threats to validity, and mitigation for the top two - What we learn if it fails - Results and conclusion fields, left blank for me to complete after running it Evidence Labelling: Mark every claim as VERIFIED, INFERRED, or UNAVAILABLE. Do not invent baseline figures, sample sizes, or conversion rates. If setting a sensible threshold requires a baseline I have not given you, say so and tell me what to go and measure first rather than inventing a number that will look authoritative later. Here is my Immediate Tactical Goal, the Intermediate Goal above it, and my constraints: [Paste your tactical goal, intermediate goal, and constraints here]

What This Library Deliberately Leaves Out

You will not find a prompt here that writes your OKRs, generates a set of KPIs for your department, or produces a strategy from a one-line brief. Those prompts would rank well and they would undermine everything the rest of this page is for.

The reason is simple. A measure is not a neutral instrument. It changes what people believe the goal is, and therefore what they do — which means choosing one is an act of leadership, not an act of drafting. Handing that choice to a model that has never met your customers produces something fluent, plausible, and quietly wrong. By the time the behaviour it drives becomes visible, it has been driving that behaviour for two quarters.

So the prompts above will help you interrogate a measure, predict how it will be gamed, and check whether it still serves the goal it came from. They will not pick it for you.

Where to Go Next

These prompts assume you already know the Evidence-Based Management framework well enough to supply real inputs. If the Key Value Areas, the goal ladder, or the experiment loop are new to you, the framework itself is the place to start — the prompts will be far more useful afterwards.

References and attribution. Evidence-Based Management, the Key Value Areas, and the goal structures referenced on this page are described in the Evidence-Based Management Guide, published by Scrum.org. The experiment canvas structure used in prompt 7 is based on the Experiment Canvas by Design A Better Business, released under a Creative Commons Attribution-ShareAlike licence. The prompts themselves are original work and are free to copy and adapt.

Ready to Lead With Evidence?

Prompts help you interrogate evidence. The course teaches you which evidence is worth gathering. PAL-EBM is a hands-on Scrum.org class for leaders accountable for outcomes — goals, measures, Key Value Areas, and the empirical loop that connects them, applied to your own organisation.

Professional Agile Leadership Evidence-Based Management certification badge