Tutoring Software With AI: What's Real and What's Hype
What can AI actually do in tutoring software?
Real AI in tutoring software does a small number of useful things well: it auto-marks structured questions, it summarises what happened in a lesson, it spots patterns in where a student keeps dropping marks, and it answers a student’s question using the actual curriculum rather than guessing. Everything else sold as “AI” is usually one of two things — a general-purpose chatbot with a badge on it, or basic automation (reminders, invoice runs) relabelled to ride the trend. The test that separates the two is simple: is the AI grounded in a real curriculum and the student’s actual work, or is it a generic model wrapped in your logo?
“AI-powered” has stopped meaning anything
Walk into any tutoring software demo in 2026 and you’ll hear “AI” within the first two minutes. It’s on every homepage, every feature list, every sales deck. Which is exactly the problem: when a word is attached to everything, it stops telling you anything.
Here’s what’s odd, though. Look closely at the established tools and most of them don’t actually ship much genuine AI at all. At the time of writing, the big operations platforms — TutorCruncher, Teachworks, TutorBird — are still fundamentally admin software; Wise.live’s parent summary is one of the few features across the category you could point at and call real AI. The category talks about AI far more than it ships it, which tells you the word is doing marketing work, not product work.
For a tutoring company owner, that’s not a small annoyance — you’re being asked to pay for it. So the useful question isn’t “does it have AI?” Almost every pitch says it does. The useful question is: what does the AI actually do, and would the product be worse without it? If the honest answer is “it writes slightly nicer reminder emails,” you’re paying for a label, not a capability.
This piece is the honest version of that conversation. Some of what gets sold as AI is genuinely useful. Some of it is a chatbot with good marketing. And a tutoring business — where the product is whether a child actually improves — is one of the worst places to be sold the hype version, because a confident, wrong answer about a maths problem doesn’t just look bad. It teaches the student something false.
What does AI genuinely do well in tutoring software?

These are real, useful, and worth paying for — because each one does something a tutor or owner can’t do as well by hand at scale.
1. Auto-marking of structured questions. When a student submits homework from a defined question bank, the system can mark it instantly and feed the result into the student’s mastery profile — no red pen, no manual entry. This is real because the questions and answers are known: the AI isn’t guessing whether an answer is right, it’s checking against a defined key and (for working) against the expected method. This is where ClassQuill’s auto-marked homework sits — tutors set homework from a curriculum-aligned question bank, students submit, results land in the mastery profile immediately. The question generation behind it took a long time to build, on leading research and specialised solving tools, precisely so the marking is accurate rather than approximately right.
2. Progress you can see — not just a lesson summary. After a session, a model can turn what happened into a clean, parent-ready summary. That’s real; summarising is exactly what language models are good at. But be honest about how much a parent actually wants another block of text. Often the sharper value isn’t a written recap at all — it’s a simple visual they can read in five seconds: a graph of their child’s marks over time, the hours they’re putting in, the accuracy climbing. That kind of view isn’t summarisation; it needs an actual learning system underneath, tracking every answer and every weakness. Most tutoring tools don’t have one — which is why they can only offer the summary, not the picture.
3. Weakness and pattern detection. Across a student’s marked work, the system can surface the topics where they keep dropping marks — patterns a tutor seeing them once a week might miss. ClassQuill’s mastery tracking and weakness detection do this topic-by-topic, and feed the pre-session view so the tutor walks in already knowing where the student struggled last time.
4. Curriculum-grounded student Q&A. A student can ask a question between sessions and get an answer anchored in their actual curriculum — not a generic internet answer. This is the one most people get wrong (more on that below), but done properly it’s real: the difference is whether the AI is constrained to the curriculum and the student’s own work, or whether it’s free-styling. ClassQuill’s between-session AI tutor is built to answer on the curriculum, not off it.
The common thread: in every real case, the AI is grounded — in a known question bank, in a real session, in a defined curriculum, in the student’s own marked work. That grounding is the whole ballgame.
What’s usually just hype?
Here’s where “AI” earns its eye-roll. None of these are useful to you as an owner, even when they demo well.
1. A general chatbot with a badge on it. If the “AI tutor” is a thin wrapper around a general-purpose model with no connection to the curriculum or the student’s work, it will be confidently wrong on exactly the things that matter. A generic model doesn’t know your state’s syllabus, doesn’t know what your tutor taught on Tuesday, and will happily invent a method that gets the student a wrong answer with full marks of confidence. On maths especially, this is dangerous, not just unimpressive.
2. “AI” stamped on basic automation. Sending a reminder when a session is 24 hours away is automation. Running an invoice batch is automation. Tagging an enquiry is automation. All useful — none of it is AI, and labelling it “AI-powered” is a tell that the marketing is doing more work than the product.
3. Anything that hallucinates on maths. This is the single clearest line in the sand. If a tool can produce a fluent, wrong solution to a maths problem and present it as correct, it is not safe to put in front of a student unsupervised. “It’s usually right” is not a standard you can run a tutoring business on — the one time it’s wrong, you’ve taught a child the wrong thing and undermined the tutor.
4. AI that doesn’t know the curriculum. A tool that can’t tell you which part of the syllabus a question maps to can’t actually track mastery, can’t generate a meaningful progress report, and can’t answer a student’s question in the context of where they are in the course. Without the curriculum underneath it, every downstream “AI” feature is decoration.
This third point isn’t abstract, and it’s worth showing rather than asserting. We ran our own maths engine head-to-head against a leading general-purpose model on 200 real VCE Maths Methods Units 3&4 Exam 2 questions. The general model — Gemini 3 Flash — scored 81%. ClassQuill, which uses advanced solving techniques and is optimised for the Australian curriculum, scored 99%. That gap is the whole point of this article in one number: a general model is impressive and still wrong often enough to teach a student the wrong method; a grounded, curriculum-optimised engine is built to not be. It’s not luck — it’s the reason the question generation and marking took so long to build on real research and specialised solving tools rather than a general chatbot with a subject prompt.
The one buyer’s test that cuts through all of it
You don’t need to understand machine learning to evaluate an AI claim. You need one question:
Is this AI grounded in a real curriculum and the student’s actual work — or is it a generic model with our logo on it?
Ask the vendor to show you, not tell you:
- “Where does the AI get its answers?” Good answer: a defined curriculum / question bank / the student’s marked work. Hype answer: vague gestures at “advanced AI” and “large language models.”
- “What happens when it doesn’t know?” Good answer: it defers, flags it, or routes to the tutor. Hype answer: it always has an answer — which means it’s guessing.
- “Show me it marking a real maths question, including the working.” If it can mark structured work against a known method, that’s real. If it can only mark multiple choice, the “AI marking” is thinner than the label suggests.
- “Can it tell me which syllabus dot-point a student is weak on?” If yes, the curriculum is genuinely underneath the product. If no, the “mastery tracking” is a session counter in a nicer chart.
- “What stops it being confidently wrong?” A vendor who’s thought about this has an answer. A vendor selling hype hasn’t, because being confidently wrong is the chatbot’s default state.
A tool that passes this test is one where the AI does real work. A tool that dodges it is selling you a badge.
Where ClassQuill sits — disclosed, not pitched
We’re not going to tell you ClassQuill’s AI is magic, because that would make us exactly the kind of vendor this piece is warning you about. Here’s the honest version.
ClassQuill’s AI is built on the grounded side of the line on purpose:
- Auto-marking runs against a curriculum-aligned question bank — known questions, known methods — not a model guessing whether an answer is right.
- Mastery tracking and weakness detection map to actual topics, so a progress report reflects real marked work, not a vibe.
- The between-session AI tutor is built to answer on the curriculum and the student’s own work, not as a free-floating chatbot.
- Pre-session intelligence comes from marked homework and practice exams already in the system — evidence, not the tutor’s memory.
That grounding is also what sets the scope. ClassQuill’s grounding is strongest in VCE Maths, and extends across STEM subjects, with English in development — because grounding is the thing that makes the AI trustworthy, we’d rather deepen it subject by subject than claim “every subject, vaguely.” The whole argument of this piece is that grounded beats broad.
The point isn’t that ClassQuill has more AI than the next tool. It’s that AI is only worth paying for when it’s grounded — and grounding is something you can test for in a 20-minute demo, with any vendor, using the five questions above. Run that test on us too.
The owner’s bottom line
You’re not buying “AI.” You’re buying answers to one question: are my students actually learning, and can I prove it to their parents? AI is worth paying for exactly to the extent that it helps you answer that — by marking real work, summarising real sessions, spotting real patterns, and answering students on the real curriculum. The moment it stops being grounded in those things, the badge is the only thing you’re paying for.
Competitors automate your operations and increasingly call it AI. The harder, more valuable thing is showing whether the learning is working — and that only counts if the intelligence behind it is grounded in a curriculum and the student’s actual work, not a chatbot with good marketing.