Danny: Yale’s AI Accounting Tutor

Overview
Danny is an AI tutor built into Yale School of Management’s accounting pre-work course. It sits beside each lesson, grounds help in the exact page a student is on, coaches reasoning before revealing answers, and turns recurring misconceptions into targeted practice.
The course prepares 350+ incoming MBA and EMBA students each year, many of them busy professionals studying between work, travel, and family. Before Danny, the course could mark answers wrong, but it could not diagnose the faulty assumption behind a mistake.
Our team redesigned the course content and modules. I led Danny end to end, from design through build.
Role
AI Tutor Product Design Lead
Responsibilities
Research synthesis, AI behavior design, UI, usability testing, full-stack build
Tools
Figma, JavaScript, Claude, Cursor
Timeline
Jan to Aug 2026
The Problem
Students weren’t stuck on content. They were stuck with no one to ask why.
The Evidence: Post-Course Survey, 875 Students
31%
looked outside the lessons for explanations
instructor videos, web searches, anything that helped
21%
called feedback the weakest part of the course
top requested fix: explain why my answer was wrong
14%
asked for more practice on core course concepts
79 of 564 answers asked for the same few concepts
User Research
Early research combined survey analysis, two learner interviews, and expert critique to define three design principles.
Who We Designed For
The Career Switcher
Composite, grounded in our MBA student interview
Learn accounting from zero and keep momentum while studying at night.
Course feedback marks answers wrong without ever explaining why.
“I want feedback that tells me why, not just whether I got it right.”
The Rusty Practitioner
Composite, grounded in our EMBA-track interview
Refresh the right concepts fast, with help that knows where he stands.
No quick way to spot which fundamentals have quietly rusted.
“There’s no professor to ask online. A tutor could remind people about these marginal concepts.”
Insight
Principle
Response
Insight
Top requested fix in the survey: explain why my answer was wrong.
Principle
Coach the reasoning, not just the answer.
Response
Danny reads reasoning and checks it against the rule.
Insight
Our expert reviewer's opening challenge: why not just ChatGPT?
Principle
Unlike a generic chatbot, help lives inside the course and cites its exact lessons.
Response
A docked panel; every claim links to the page it came from.
Insight
Both interviewees: the hard part is concepts that quietly stay broken.
Principle
Remember what keeps breaking for each student.
Response
A misconception model with receipts, feeding targeted practice.
Product System
The solution: an AI tutor woven through the course, aware of the lesson page and the student’s progress. Five surfaces in one window.





Information Architecture
How Danny Thinks
Before any screens, the tutor’s behavior had to be designed: from a stuck moment to targeted practice.






Three Use Cases
One student, one misconception, three moments where Danny earns its place.
The three moments follow one student and one misconception. She asks about the concept, misapplies it in her own reasoning, and Danny turns the recurring pattern into practice. HoverTap each to watch it play out; the deep dives below unpack each one.
A Question Mid-Reading
“I'm confused now, but leaving the page breaks my flow.” Danny answers beside the lesson and cites the exact course page.
Stuck on a Judgment Call
“I can defend my answer, so why is it wrong?” Danny checks the reasoning against the rule, names the trap, and hands the deeper why to a follow-up chat.
The Mistake That Keeps Coming Back
“The same mistake keeps happening and nothing sees the pattern.” Danny tags mistakes across lessons and turns them into targeted practice.
Deep Dive: Use Case 1
A Question Mid-Reading
A student still learning the subject can’t tell a right answer from a confident wrong one.
The Design Challenge
Let students ask mid-lesson without breaking their flow, and back every answer with a course source they can click to check.
Iteration
v1 opened in its own tab because the lessons lived on another platform: Danny couldn’t see the page a student was reading, and students flipped between tabs to chase what an answer pointed at. Pilot students asked, unprompted, for the tutor beside the content, so I rebuilt the lessons into the Study Window and docked Danny next to them.


Voice went the same way: students mostly study in quiet places where talking out loud isn’t an option, so calls became a button inside the composer, not a mode.


Result
Answers worth keeping become notes: concept cards with claim, takeaway, and grounding. Course-specificity became the most praised quality in post-study feedback.
87%
of pilot students used no outside AI or materials
Danny was effectively the whole support surface
0
answers drifted off the course material
Danny declined off-topic questions and flagged concepts beyond the modules
Deep Dive: Use Case 2
Stuck on a Judgment Call
The course can mark an answer wrong, but it can’t see the wrong idea behind it.
The Design Challenge
Point out the wrong idea without giving away the answer, because once Danny tells, the student stops thinking.
Guardrail
Danny never grades. The course’s own checks decide whether a number is right; Danny steps in after the check, coaching the thinking behind the answer.
Iteration
The course’s authored checks stayed; they grade every definite answer. But pre-written feedback can’t say where a specific miss came from, and pilot transcripts showed those questions piling up in the hardest lessons. So I added a second layer: Danny reads the answer itself, a number or the student’s own reasoning.


Result
In the A/B pilot, students with Danny made fewer mistakes and got more answers right on the first try.
“The results are reliable… helped me answer the questions with more confidence.”
Deep Dive: Use Case 3
The Mistake That Keeps Coming Back
If Danny claims a student keeps making the same mistake, it has to show the proof.
The Design Challenge
Every weak spot Danny names has to link back to the exact moments it happened, so the student can check the claim themselves.
Iteration
v1 tracked completion: a percent ring per page, with checklists of what you’d opened and finished. It told students where they’d been, but nothing about whether the ideas stuck. v2 tracks understanding instead: each concept sits at one of four levels, Getting There, Practiced, Solid, Mastered, and moves only on what the student actually got right and wrong. Page completion didn’t disappear; it moved into the page index, counting pages done per concept.



Result
Students can open any level and see why it’s there. Progress reads as evidence, not a guess. From the pilot, on the practice feature: “It was nice to see my overall progress.”


The Design System
The moments read as one product because every surface speaks one language.
Color
Green and red: outcomes only, never decoration.
Typography
Titles, 21/600
Body reads at 16/400.
Buttons sit at 16/550.
Tags and labels, 14/600.
Meta at 12, muted.
SF Pro · Segoe UI · Roboto, the native stack.
Emphasis is weight, never italics.
Spacing
A 4px grid: inset 20, gaps 24.
Avatar
Glass finish and glow live in the artwork.
Buttons
One primary per view; an arrow means it navigates.
Inputs
Cards
Tokens
Patterns
Course left, Danny docked right: 920 / 520 on a 1440 screen; the dock drags 340 to 680.
Read, get checked, practice.
Accessibility
Every text pair holds AA, 4.5:1 or better.
Danny on Mobile
EMBA studying happens on trains and between meetings. The system follows onto mobile: the same brain, in the phone’s native form.

Danny docks at the bottom, not the side.

The diagnosis and the drills, one column.

Notes ride along for the commute.
Outcome
The pilot pointed the right direction. I’m careful about how far it points.
Clean completion
12% aheadFirst-try correct
11% aheadAvg. mistakes
41% fewerThirty-one students, fifteen with Danny and sixteen without. Directionally consistent, and honestly framed: a pilot without significance testing, where usefulness ratings tracked how much each student leaned on Danny in the short study window. The caveats shaped the next questions more than the wins did.
In Their Words
“It worked great because it was so course specific.”
Takeaways
What Improved
- Help moved inside the course, where the struggle happens
- Answer-first tutoring became method-first coaching
- Misconceptions became visible, auditable evidence
- Notes, practice, and progress share one learner model
What I’d Test Next
- Does method-first coaching improve retention weeks later?
- Does misconception-based practice reduce repeated errors?
- Do students trust Danny appropriately, without over-relying on it?
The Clearest Signal
The school is exploring Danny for courses in other subjects. Accounting was the hardest first test: a formal system where a wrong idea can hide inside a right-looking number.
Technical Tradeoffs
I built the product I could defend, and I know where the design ends and the engineering begins.
Behavior as Policy, Before the Model
What Danny cites, asks, and declines was authored as policy first, so the model has a spec to meet.
Grounded, Not Trained
Danny grounds each reply in the real lesson pages it retrieves, so answers cite exact sources and stay in scope.
From the Course, Already Signed In
Danny plugs into the course site as a standard LTI 1.3 tool: one click opens it in place, no second login, no password.
A System, Not a Subject
Swap in another course's lessons and the same coaching and learner model carry over: accounting was the starting point.
