How I decided where to put AI in an educational product (and where not to)
I built Examen UNAM, an exam-prep platform for Mexico’s largest public university admission exam. Students sit a timed simulation of the real test, and afterwards the product tells them where they stand and what to do next.
That “what to do next” is the part everyone assumes is AI. Most of it isn’t, on purpose.
What follows is the actual map of where the model runs, where it is deliberately absent, and the reasoning behind each placement. Every decision here shipped to production. Every one of them cost something, and I have tried to name that cost honestly, because a decision you can only defend by its upside is a decision you haven’t finished making.
The rule everything collapses into
The screen shows the what. The model explains the why. Anything that can be computed is computed, and the model is only invited in where reading matters more than counting.
That was not the starting point. The first prompts asked the model to do everything, including redrawing the subject-by-subject table that the results page already renders underneath it. It worked, in the sense that it produced words. It was also the most expensive way imaginable to print a number the browser already had.
The rest of this piece is the list of things that got moved out of the model, and the few that stayed in.
2 points below the historical admission cutoff
Computer Science · representative campus
Your result is near the historical cutoff, but one attempt does not predict admission. Physics and Chemistry contain the largest recoverable gaps; start there, then use Spanish to protect the points you already earn consistently.
- Physics55%
- Chemistry68%
- Spanish82%
Neutral by design: the example neither predicts admission nor turns a near miss into failure. It states the current result and identifies the next useful action.
The map
- 01
Post-exam analysisStrong model · once per attempt · cached
- 02
Post-practice advisoryCheaper model · short · cached
- 03
Curated content supportOffline: model drafts → human edits
Runtime: code selects approved facts → model explains
- Score vs. cutoff
- Subject & topic breakdowns
- Action-plan ordering
- Weak-subject buttons
- Progress across attempts
- Free demo analysis
- Failed-question explanations
- Study calendar DEFERRED · NOT SHIPPED
Boundary: numbers, sorting and thresholds stay in code. Reading, synthesis and explanation earn a model call.
Where the model runs, in three places and only three:
- Post-exam analysis. One generation per exam attempt, cached forever. This is the product’s central AI feature and the one that justifies the paid plans.
- Post-practice advisory. One short generation per practice session, on a cheaper model, cached the same way.
- Offline content generation, curated by a human. Review cards are generated in batch and edited before any student sees them. At runtime the model only selects and prioritizes two or three topics and explains why it picked them. It never writes facts live.
Where the model is deliberately absent:
- The score against the admission cutoff. Pure data: the student’s score crossed with the official cutoff for their program. This is the product’s core promise and it never touches a model.
- Subject and topic breakdowns. Computed and drawn by the UI, ordered worst to best. The model receives them as input and is forbidden from reproducing them.
- The action plan steps. A deterministic function decides what to practice first and in what order. The model, on the higher tiers, writes one paragraph per step. On the entry tier the same steps ship with fixed descriptions and no model text at all.
- “Practice your weakest subjects” buttons. A threshold and a sort, not a recommendation engine.
- Progress across attempts. Computed in code with explicit thresholds. The model receives three already-resolved lists: regressed, improved, persistent.
- The free demo’s analysis. Fully hardcoded branching on effort-to-target. It produces a narrative, a study cadence and a plan pitch without a single model call.
- Explanations of failed questions. Stored with the question bank, curated, never generated on the fly.
- The study calendar. It turned out to be a different product, not a longer version of the per-attempt analysis. It does not exist yet, and the pricing page no longer promises it.
The pattern: anything that is a number, a sort or a threshold is code. The model gets the three things only reading can produce: what a specific error says (“the mistake was a sign error, not a method error”), gaps that cross subject boundaries (proportionality failing in both physics and chemistry, invisible when topics are grouped by subject), and the weight argument (6 of 22 in math and 2 of 10 in chemistry look identical as percentages, but one is 16 recoverable points and the other is 8).
The trade-off: this map makes the product less impressive to demo. A competitor can say “our AI writes your whole study plan”; I have to say “our AI writes the paragraph that explains a plan our code built.” The narrow version is cheaper, it cannot hallucinate your score, and it degrades into something still useful when the provider is down. I took the weaker sentence in exchange for the stronger product.
| Alternative considered | Decision taken | Benefit | Accepted cost |
|---|---|---|---|
| One model for every interaction | Strong model for exams; cheaper model for practice | Sustainable cost at real usage volume | Uneven model capability across surfaces |
| Regenerate on demand | Generate once and cache permanently | Predictable cost and instant revisits | No “try again” button after a weak result |
| Interpret every score change | Stay silent below the evidence threshold | Avoid confident stories about statistical noise | Silence can feel like a missing feature |
| Keep the promised study calendar | Remove it from pricing until it exists | Product and marketing cannot drift apart | A less ambitious sales story |
The economic decisions
The tier did not choose the model. The interaction count did.
Exams are few and expensive to take: five, ten or twenty a semester depending on the plan, three hours each. Practice sessions are many, and unlimited on the top tier. Those two surfaces got different models on purpose.
Spend the reasoning budget where the interaction is rare and consequential.
The exam analysis runs on the stronger model for every plan, including the cheapest. The practice advisory runs on the cheapest model for every plan, including the most expensive. The arithmetic that decided it: with the stronger model, worst-case practice usage on the top tier would consume roughly a quarter of what that plan costs. With the cheaper model it stays in the low single digits.
The trade-off: the customer paying the most gets the cheap model on the surface they use most often. That reads backwards until you notice the alternative is a plan that loses money precisely on its best customers. The advisory is 120 to 220 words of “here’s the pattern in these ten questions,” a job the cheap model does well. The three-hour exam analysis is where the reasoning has to hold, and that is where the budget went.
Depth follows the price where volume is capped, and the content where it isn’t
On the exam, depth scales with the plan: 2.5x deeper at the middle tier, 5x at the top, in words and in section count. Those ratios on the pricing page are the literal word budgets in the config.
On practice, depth scales with the size of the session, not with the plan. If the top tier multiplied words-per-session by unlimited sessions, cost would be unbounded. All three tiers get the same promise, advice at the end of every session, and the cost per session follows the content.
The trade-off: two different rules for two surfaces is harder to explain on a pricing page than one rule everywhere. I accepted the explanation cost because the alternative was an uncapped liability sold at a flat monthly price.
The pricing page can only promise what a flag actually gates
Everything that differentiates the three tiers lives in one configuration object: word budget, section count, how many error samples the model sees, whether it writes a paragraph per plan step, whether it gets a timing-strategy section. The pricing matrix derives its rows from that same object, and tests fail if the two drift apart.
Stated as a rule: a line on the pricing page can only differ between tiers if the code actually gates it. That rule killed a row reading “Personalized AI study plan (coming soon)” on the top tier. It lied twice: it promised a product that did not exist, and it implied the middle tier didn’t get something the middle tier already received.
The trade-off: you lose the ability to put aspirational rows on the pricing page, which is a real marketing tool and not always a dishonest one. What you get back is a page that cannot rot. Nobody has to remember to update it.
Generate once, cache forever: why validation happens before the write
Each analysis is generated once per attempt and never regenerated. Regenerating would be money spent to replace something the student already read. Two consequences follow:
- The endpoint is idempotent and locked. If an analysis exists it is returned without calling the provider, and a generation in flight blocks a second one.
- Validation happens before writing, never after reading. A malformed payload gets one retry with the validation error fed back to the model; if it fails again the request fails and nothing is stored. Persisting a bad payload means persisting it forever.
The trade-off: a student who gets a weak analysis is stuck with it, and I gave up the easy “regenerate” button. Caching forever is only defensible if the write path is strict, so the strictness is the price of the caching, not a separate virtue.
Measure the budget instead of assuming it
Output budgets were originally set by hand, assuming roughly 2 tokens per word. Measured against the actual model in Spanish, it is about 2.8, because of accents and long words. The old budgets truncated the top tier mid-sentence and left the middle tier sitting at 96% of its ceiling.
A related trap: the newer model reasons by default, and its reasoning is charged against the same output budget. Reasoning is disabled deliberately, because budgets calibrated for visible text get eaten by hidden text. If it is ever turned on, the budget rises in the same change.
The trade-off: turning reasoning off costs quality on the hardest analyses. I would rather have a slightly less clever analysis that finishes its last sentence than a better one that stops mid-paragraph because invisible tokens ate the ceiling.
Structured output: pay 20% more tokens to stop parse failures
Moving the analysis from markdown to a JSON contract costs roughly 4% in envelope characters, and up to 17% when the model escapes accented characters. The budget went up 20% to cover it.
The trade-off: this is the clearest one in the project. A truncated markdown document reads badly; a truncated object does not parse and the entire analysis fails. Going structured converts a soft failure into a hard one. That is worth it for the reliability of the contract with the UI, but only if you also loosen the budget, and the instinct when costs go up is to tighten it.
The statistical decisions, or: a well-placed “no” is a capability
This is the section I would point a client to first, because it is the part that is invisible when it works.
A 120-question exam over ~45 topics is two or three questions per topic
At that sample size, a topic going from 1-of-1 to 0-of-1 is a coin flip, not a regression. Worse, the questions change between attempts, so “last time they got the two easy ones, this time the two hard ones” looks exactly like forgetting. Feed that raw comparison into a prompt and the model will write a very convincing narrative about pure chance.
The model never decides whether evidence is sufficient. Code resolves that first; the model only explains signals that passed the gate.
So:
- Subjects (10 to 22 questions each) can be compared directly. Topics can only be compared by accumulating several attempts.
- The cross-attempt comparison is computed in code with explicit thresholds: a minimum of three accumulated questions in history and two in the current exam, mastery at 70%, failing at 50%, five items per list maximum. The model gets three pre-filtered lists and no raw history.
- If there is no signal, the section is not sent at all. The system prompt tells the model that an empty list means lack of evidence, not that everything is fine.
- A “regression” is almost never forgetting, and the prompt says so. It usually means the topic was never solid.
The trade-off: the product stays quiet exactly when students most want to hear that they are improving. Silence feels like a missing feature, and it is the single hardest decision here to defend in a demo. But a study tool that tells you to drop a topic you never actually lost has spent the student’s scarcest resource, time before the exam, on noise.
Constant prompt cost regardless of how long the history is
The first design capped history at two previous attempts to save tokens. Once the comparison moved into code, prompt cost stopped growing with the number of exams, so the window went up to five attempts over 180 days. Per-subject trajectory is one line per subject, 4/22 → 5/22 → 6/22, not one table per attempt, because five tables is fifty rows and the model would recite them back.
Time is sanitized before the model ever sees it
- An exam date already in the past is treated as unknown before it reaches the prompt. Talking to a student about deadlines for an exam that already happened is worse than not mentioning deadlines at all.
- Days-to-exam are computed in the student’s timezone, not the server’s. The server runs in UTC, where 7pm in Mexico City is already tomorrow.
- History age is measured against the analyzed attempt’s finish time, not against now, because the analysis is cached forever and would otherwise drift into telling a student that “recent” data is two years old.
One gap I know about and chose not to close
The per-topic breakdown inside practice still turns single-question samples into verdicts: a 0-of-1 in red, seven 1-of-1s in green. It is the same statistical mistake I spent the previous section designing out of the exam analysis. I found it, root-caused it (a component built for subject-level granularity, reused at topic level), and left it.
The trade-off: shipping a known flaw is a real cost, and naming it publicly is a second one. It stayed because practice is a low-stakes surface where the student sees the question they just missed, and fixing it properly meant reopening a component used in four places. Knowing where your product lies to itself, and choosing when that gets fixed, is a different skill from not having bugs.
The contract between the model and the interface
A forced tool call instead of “please return JSON”
The analysis contains LaTeX. Inside a JSON string every backslash doubles, which is precisely where a model hand-writing JSON slips, and the result is a parse failure. With a forced tool call the SDK does the serializing and escaping stops being the model’s problem. One schema is the source for both runtime validation and the tool’s input schema.
The model names targets, never labels
The model writes target: "mathematics" and prose. It never writes the subject’s display name, and never the “+21 possible points” chip. The list of valid targets comes from the same function that renders the plan page. If the prompt built its own list, sooner or later a student would read a physics paragraph under a chemistry heading. An invented target is silently discarded, and a step with no paragraph falls back to deterministic text.
The trade-off: the prose is slightly stiffer, because the model is writing around nouns it isn’t allowed to print. In exchange, a wrong label, the kind of error that destroys trust instantly and permanently, is not reachable from the model’s output.
Ban duplication, then explicitly ask for interpretation
Telling the model “do not repeat the subject list the screen already shows” made it stop mentioning subjects altogether. The fix was not to remove the ban but to add its other half: your job is to interpret the data, naming subjects and topics whenever the explanation needs them; what is forbidden is redrawing the list. Prohibitions alone produce evasive output. You have to say what the space you cleared is for.
When output looks broken, check the instructions before the parser
Reported bug: formulas not rendering. Verified: the math pipeline was fine. The prompts said “no lists, no headings” and never specified which LaTeX delimiters to use, so the model reached for the ones the renderer didn’t understand. The bug was in the prompt, two layers away from where it showed up.
What the model is allowed to say about a student’s future
These are product decisions more than AI decisions, but they set the voice the model inherits, and they are the calls a founder is actually buying when they hire someone like me.
- Equivalence, never projection. A demo score is shown as an equivalent on the real 120-point scale, never as a prediction. Same number, different word. Equivalence is a scale conversion: true and verifiable. Projection is a forecast carrying roughly ±21 points of uncertainty when it comes from 27 questions. No future tense anywhere: no “you will get,” no “you’re on track for.”
- The private page tells the whole truth; the public card tells only what a student would want to send. The results page says “you are 12 points short.” The shareable card omits it.
- Good news needs a margin. On the demo, “you clear the cutoff” only appears with 10 or more points of headroom, because a good-news-only message is exactly the one that gets contradicted on exam day.
- Don’t sell when the goal is already met. If the demo score clears the cutoff comfortably, the plan pitch is suppressed and the call to action becomes a plain registration.
- The reference cutoff is not the most recent one. One cycle was inflated, the control that followed was deflated. A single function picks the reference cycle for every surface, excludes non-comparable ones by explicit list plus shape detection, and always labels the number with the cycle it came from. Retiring those exclusions is scheduled, not automatic.
- No prediction of the cutoff. Ever. Not in the product, not in the model’s output, not in the marketing.
The trade-off: every one of these leaves conversion on the table. Suppressing the pitch when a student is already passing is, straightforwardly, a decision not to sell to someone who might have bought. I’d make it again: this is a product students use in the highest-stakes month of their academic lives, and the first time it tells someone they’re fine when they aren’t, it is finished.
What this adds up to
None of these are exotic decisions. They are the ordinary ones, made in order, with the numbers in front of me: what does this interaction cost at the volume it will actually run at, what could a simpler system do instead, and what does waiting really cost.
That’s the framework, and it’s published in full at apat.io/framework. This is what it looks like when it runs for a year against a real product.