The most useful number in the AI workload debate is not a promise from a vendor. It is 25.3 minutes.
That is the average weekly difference in lesson and resource preparation time between two groups of Key Stage 3 science teachers in an Education Endowment Foundation trial. Teachers using ChatGPT with a guide spent 56.2 minutes on the measured preparation, compared with 81.5 minutes for teachers asked not to use generative AI. The EEF reports a 31% reduction and says an expert panel found no apparent difference in the quality of the lesson resources.
Those findings were published in the trial's evaluation in 2024. They are worth revisiting as schools write AI policies this term, but this is not a new September 2026 result.
What did the teachers actually do?
The trial involved 259 Year 7 and 8 science teachers in 68 state-funded secondary schools in England. Teachers in the ChatGPT group had five weeks to familiarise themselves with a guide and practise before preparation time was recorded over the following five weeks.
They were not handing entire schemes of work to a chatbot. The EEF says teachers commonly used it for one or two activities in a lesson: finding ideas, making questions or drafting a quiz. That limited use is part of the finding. It is also a far more realistic starting point than the sales pitch in which an AI writes a week's teaching before the kettle boils.
The trial was about a specific subject, year group and task. It was not a test of whether children learned more science. Nor did it show that an inexperienced teacher would save the same amount of time, or that an unchecked generated explanation would be safe to put in front of pupils.
The work moved; it did not disappear
Someone still had to choose what the class needed, check the science, match the questions to the curriculum and decide which suggestions to discard. That is teacher work. The tool may have made one part of it faster.
There is also an up-front cost. The five-week learning period happened before the measured preparation period. A school that buys access on Monday and asks for proof of time savings by Friday has copied the product but missed a condition of the study.
Even the resource-quality result has a boundary. A panel reviewed the materials without knowing which group made them. That is useful. It cannot tell us how well the lesson was taught, what pupils retained, or whether AI changed the teacher's understanding of the topic.
A better department experiment
Pick a task you repeat every week, such as writing retrieval questions for Year 8 science. For a short period, record how long it takes without AI. Then let willing teachers try an approved tool, give them time to learn it, and record the time again. Keep the subject review: are the questions accurate, appropriately difficult and tied to what the class has actually studied?
If the time falls and the questions hold up, you have a local reason to continue. If the teacher spends the saved minutes correcting plausible errors, the net gain may be smaller. If the tool generates fluent nonsense about a topic the teacher does not know well, stop there.
This is the AI literacy task for staff. Know what the model is good at, spot where it is guessing, and keep the decision that matters with the teacher. The EEF trial gives schools a credible reason to investigate that workflow. It does not give them permission to skip the investigation.
Source: Education Endowment Foundation, ChatGPT in lesson preparation: Teacher Choices trial. The evaluation report was uploaded in December 2024; the project page now labels the project complete.