AI under scrutiny: This is how artificial intelligence was evaluated on the French baccalaureate philosophy exam.

  • ChatGPT was submitted to the French Baccalaureate philosophy exam, and the response created was evaluated by a teacher and several AIs.
  • Human evaluation detected conceptual errors, a lack of depth, and question reformulation, scoring the essay 8/20.
  • The AI ​​tools awarded much higher grades, failing to acknowledge the underlying flaws noted by the teacher.
  • The experiment highlights the limits of AI in philosophical reasoning and the gap with human judgment.

Philosophy exam with AI

In recent days, a curious An educational experiment in France has tested the real capabilities of artificial intelligence. to face the philosophy exams for high school, the well-known Baccalaureate. The trigger was an open question: Is the truth always convincing? A typical question on university entrance exams, one that measures students' argumentative maturity right at the end of secondary school.

France 3 Hauts-de-France, a public broadcaster, decided to commission ChatGPT to write an essay. As if it were a student aiming for the highest grade. The goal? To test how far an AI can overcome teacher filtering and automatic grading tools.

The proposal to ChatGPT and the teaching criteria

To mimic the real situation, the AI ​​was provided with a Detailed prompt: should adopt the style and structure of a senior student, organize the text into an introduction, development, and conclusion, and address all the nuances of the topic.
When the generated response was presented, at first glance the writing seemed academically correct: fluent sentences, absence of spelling errors, and a clear structure. However, the initial impression crumbled upon closer examination.

The philosophy professor in charge of correcting the essay gave it an 8 out of 20.. Why such a low score? Mainly because detected a lack of depth in the arguments, a lack of examples, and, above all, an unexpected twist in the way the question was worded: the AI ​​went from answering "Is the truth always convincing?" to asking "Is the truth sufficient to convince?" For the teacher, this change showed that the system hadn't fully understood the original instruction, a significant error in this type of test.

Another negative aspect pointed out was ChatGPT's tendency to repeating standard formulas and avoiding personal reflection, which made the result too superficial compared to what is expected of a well-prepared student.

manuscripts-0
Related article:
New findings and technologies in the study of historical manuscripts

Other AI systems and differences in criteria

The exam was not limited to the teacher's opinion. The text generated by ChatGPT was also evaluated by different AIs., including Gemini, Perplexity, DeepSeek and Copilot. All of them agreed on give much higher scores: between 15 and 19,5 out of 20.

What accounts for this striking difference? The automated platforms emphasized the essay's good formal structure and surface coherence, but none detected the key error in understanding the topic nor the lack of argumentative precision identified by the teacher. Moreover, ChatGPT herself gave herself a 19,5/20 rating, showing little self-criticism.

For the teacher, all this confirms that AIs can meet mechanical requirements well —orderly writing, connectors, basic examples—, but they are unable to go deeper, nuance or grasp the conceptual and philosophical nuances required in these exercises.

Reflecting the current limits of AI in education

This case has served to put on the table The current limits of artificial intelligence when it comes to philosophical reflection and critical analysis. Although programs like ChatGPT handle formal aspects very well and produce seemingly convincing texts, The ability to argue, question or contribute one's own points of view is still far inferior to that of the real student.

The essay is valued, beyond the well-connected words, did not demonstrate original reasoning or respond accurately to requestsThe teacher commented that A student would have thought about everything that was missing and would have done much better..

On the other hand, the fact that AI tools give such favorable ratings to texts generated by other AI highlights the existence of biases in automatic evaluation systems, which prioritize form over substance and are less demanding in conceptual analysis.

the lights of february
Related article:
The lights of February: Joana Marcús

Add as preferred source in Google