AI and open-ended answers: what works, what doesn't yet

Multiple-choice questions naturally lend themselves to automated evaluation. Open-ended, argumentative answers are a far more complex challenge, and it's worth understanding why before applying AI broadly across every type of assessment.
What works reasonably well
AI can effectively evaluate the presence of key concepts in a response, the general structure of an argument, and certain surface elements like clarity or basic logical coherence. On these dimensions, first-pass assistance can be fairly reliable, especially when criteria are well defined ahead of time.
What remains harder
Evaluating the actual quality of an original argument, appreciating a creative perspective that steps outside the expected framework, or judging the depth of personal reflection are tasks where automated judgment stays far less reliable. A text can be well-structured and grammatically correct while completely lacking substance — or the opposite: an awkwardly written response can contain a particularly sharp idea.
Why this distinction matters
Understanding this limit helps decide where AI can be useful as a first pass, and where a full human read remains essential from the start. Using AI on tasks where it's unreliable creates more grading work than time saved, defeating the original purpose.
A technology to keep watching
AI systems' capabilities keep evolving, and it's plausible that handling of open-ended answers will improve over time. That said, it's wise to base current decisions on today's actual capabilities, not on unguaranteed promises of future progress from technology vendors.
A realistic combination
The best approach today generally combines AI-assisted evaluation for objective criteria with a full human read for assessing argumentative substance. It's a division of labour, not a replacement, and that division is what delivers the best results in practice.
An example that illustrates the limit
Two students might defend the same thesis in an essay, one with a conventional, well-structured argument, the other with an original angle that departs from the expected outline. An automated system will often default to favouring the first, while a human reader can recognize the value of the second.
One last nuance
This limitation also varies by subject. Open-ended answers in math or science, where there's often a single valid approach, are generally easier for a system to evaluate than open-ended answers in literature or philosophy, where the range of valid approaches is much wider.
Understanding where this limit currently sits makes it possible to make informed decisions now, without waiting for a perfect solution that might never arrive on the timeline an eager institution would prefer.
Honestly acknowledging this limit, rather than ignoring it, lets an institution make technology choices better aligned with the reality of the subjects it teaches, rather than with a vendor's general promises.
Sharing this kind of subject-specific insight with colleagues teaching the same course helps the whole team calibrate their expectations together, rather than each person arriving at a different comfort level through trial and error alone.
That kind of shared calibration tends to produce more consistent outcomes across a whole department.
Being upfront about this variation, rather than promising uniform results across every subject, builds more durable trust with teaching staff.