All articles

The Research ClassQuill Is Built On — And What It Means for Your Tutoring Business

By Brandon Collis 7 min read
ai-tutoring-researchtutoring-management-softwareevidence-based-tutoringclassquillb2b
The Research ClassQuill Is Built On — And What It Means for Your Tutoring Business

What research actually backs ClassQuill’s AI features?

Three independent, peer-reviewed studies: a Stanford randomised controlled trial (arXiv:2410.03017) showing an AI co-pilot lifts student mastery by 4 percentage points overall and 9 points for the weakest tutors; McGill research (arXiv:2505.11899) on generating practice questions at the right cognitive depth rather than just “more of the same”; and CMU/BEA work on automatically scoring session quality with the reliability of a human rater. ClassQuill’s AI features are built directly from these findings, not from an internal hunch about what “personalisation” should look like.

When a new platform tells you their AI “personalises learning” or “adapts to each student,” what they usually mean is: we have an algorithm, we think it helps, we haven’t actually measured it.

I spent a year reading the published research on AI-assisted tutoring before building ClassQuill. Not to put logos on a landing page — to understand what actually moves the needle when a tutor sits with a student. The findings are more specific, and more useful, than most EdTech marketing lets on. So here’s what the science actually says, what we built from it, and why it should change how you think about running sessions.


Does an AI suggestion during a live session actually help? (Finding 1)

In 2025, Stanford researchers Wang et al. ran the first properly controlled trial of AI-assisted live tutoring. Nine hundred tutors, 1,800 students, real sessions. Half got access to a real-time AI sidebar suggesting moves during sessions — things like “try asking a question here” or “give a hint rather than the answer.” Half didn’t.

The result: tutors with the AI sidebar improved student mastery by 4 percentage points on average. For the tutors who started with the lowest effectiveness ratings, the gain was 9 percentage points.

The cost: approximately $20 per tutor per year.

There are two things worth unpacking here. First, the gain was not from the AI doing the tutoring — the tutor was still in the room, making every decision. The AI just prompted better moves at the moments that mattered. Second, the gain was largest for the weakest tutors. That’s not an accident. The AI was effectively giving them access to what good tutors already do instinctively — asking questions instead of explaining, giving hints instead of answers, staying Socratic when the student is stuck.

This is built into ClassQuill as the TutorCopilot panel. In every live session, your tutors can see AI-generated suggestions in real time. It’s not a co-pilot taking the wheel. It’s a prompt at the moment a tutor might otherwise jump to the explanation.

What this means for you: If you have tutors who are effective but inconsistent — or newer tutors who are still developing their instincts — this is the most direct lever research has identified. You don’t need to replace them. You just need to give them better prompts at the right moment.


Finding 2: The moves your tutors make are measurable — and predict whether students learn

The Stanford paper also introduced a taxonomy of tutor moves that the AI uses to classify what a tutor is doing. Seven strategies, ranging from high-value to low:

Move What it looks like Pedagogical value
Ask a question “What do you think happens if you substitute here?” High
Provide a hint “Think about what you know about the reciprocal rule” High
Provide a similar problem “Try this simpler version first” High
Explain a concept Full explanation of the method Neutral — context-dependent
Affirm a correct answer “Yes, that’s right” Low if that’s all you do
Encourage the student “Keep going, you’re close” Low pedagogical value alone

The research is blunt: sessions dominated by “ask a question” and “provide a hint” produce better outcomes than sessions dominated by “explain a concept” — even when the explanation is correct. The reason is that explanation is passive reception. A question forces the student to retrieve and apply, which is what actually builds durable memory.

This isn’t new to experienced tutors. What’s new is that it’s now measurable per session. ClassQuill logs which moves are being used across sessions, and the post-session auditing pipeline (coming in the next build) will surface which tutors are predominantly Socratic and which are defaulting to explanation mode. That turns what was previously intuition-based coaching into something you can act on with data.


Are all practice questions equally useful? (Finding 3)

McGill researchers Yu, Krantz, and Lobczowski (2025) spent a year testing whether LLMs could reliably generate maths questions at controlled cognitive depth levels. They used Webb’s Depth of Knowledge (DOK) framework, which separates questions into four tiers:

  • DOK 1: Recall facts and definitions
  • DOK 2: Apply a known method to a routine problem
  • DOK 3: Reason strategically in a non-routine scenario
  • DOK 4: Connect multiple concepts across a complex problem

Their finding: when the question generator was given only a topic name (no curriculum context, no DOK target), it produced questions clustering around DOK 2.4 on average — mostly routine, occasionally higher. When they added both a DOK level target AND RAG (retrieval of the actual curriculum content), correctness went from 0.60 to 1.00, and appropriateness to 0.92. Lexical diversity — meaning the questions weren’t just paraphrases of each other — scored 0.92.

The practical translation: a question bank that just pulls “more questions like this” is generating DOK 1–2 practice for students who need DOK 3–4 to be ready for exams. The VCE Mathematics examinations test DOK 3–4 almost exclusively.

ClassQuill’s question generator targets DOK levels explicitly, grounded in the actual VCE curriculum content, not generic maths. When a student has a gap at a particular topic, the practice generated is calibrated to where they are — not a generic harder question.

What this means for you: Practice quantity is not the constraint. Most students already do more questions than they need. The constraint is practice quality — specifically, practice that forces the type of reasoning the exam actually tests. Targeting DOK level is how you close that gap.


What’s coming: post-session quality scoring

Researchers at CMU (Thomas et al., 2025) developed a method for scoring the pedagogical quality of tutoring sessions using LLMs — not by asking for a 1–10 rating (which produces meaningless averages), but by asking a structured two-question pattern: did this situation occur? If yes: was it handled well?

Applied to live tutoring transcripts, their method detected praise events with 94–98% accuracy and error-handling events with 82–88% accuracy, matching the reliability of human raters.

Separately, the BEA 2025 benchmark (Kochmar et al.) established four dimensions for evaluating tutoring quality: whether the tutor identified the student’s mistake, whether they located it precisely, whether they guided the student toward the answer (rather than giving it), and whether the feedback was actionable. The hardest dimension — and the most important — is “providing guidance,” which is precisely the distinction between Socratic tutoring and explanation.

ClassQuill’s next phase combines both: after each session, an automated pipeline scores the transcript across all four dimensions and surfaces the results in the org admin dashboard. Tutors don’t see their individual session scores (not appropriate without a coaching context). Org owners and managers do — as per-tutor quality trends over time, flagged sessions below threshold, and a breakdown by dimension.

This is what actually makes quality visible at scale. Not “how many sessions did each tutor run” but “are your tutors guiding students or just explaining at them?”


The rule underneath all of this

Notice what none of these research findings are saying: “replace the tutor with AI.” Every finding above is about making the human tutor more effective — better prompts in the moment, measurable patterns over time, questions calibrated to the right cognitive level.

The AI in ClassQuill does the high-volume, low-judgement work. Suggesting moves. Generating practice. Scoring transcripts. The tutor does the part that requires a human who knows the student and the subject.

If you’re evaluating tutoring software and the pitch is “our AI teaches the student,” ask to see the RCT. The research doesn’t support that claim. What it does support is AI that makes skilled tutors more consistent — and that’s what we built.


ClassQuill is tutoring management software for tutoring companies and organisations — lesson delivery, parent updates, question generation, and now session quality auditing in one place. See how it works →