School Report Narrative Engine
A Concept Brief
First Edition — 12 March 2026 · Prepared by Greg Williams · steko.co.nz/thinking
Editorial notes — 30 & 31 July 2026
In brief: why this research exists; what two practising teachers said when asked to test its assumptions against their working reality; why the concept is not going to market as a teacher-facing product; one correction from the March analysis, disclosed in full; and the reader survey now gathering evidence for the second edition.
Read the full editorial notes
30 July 2026
This research began with years of hearing, first-hand, what reporting season costs the teachers in my life. The after-hours toll is the reason the study exists.
Since the concept brief was completed in March, I have shared it with two practising teachers (one primary, one secondary) and asked them to test its assumptions against their working reality. Their answers confirmed more than I expected. The secondary teacher described end-of-year report comments assembled from a faculty comment bank, personalised only when time allowed, which was rare, written in late nights and weekends with no allocated time. The primary teacher described a twenty-to-thirty-hour reporting cycle that begins in the school holidays, then five to ten further hours reworking comments through the review chain to fit character-limited boxes. Both described the same information vacuum when a child arrives from another school: records that follow weeks later, or never.
Neither rejected the idea of AI in reporting. Their concerns were precise and reasonable: output should sound like the teacher who wrote it, it must be checked, and privacy must hold. One volunteered, unprompted, that a longitudinal record following a child through their schooling could be genuinely transformative.
That feedback has shaped what happens next. I am not taking this concept to market as a teacher-facing product at present, for three reasons. First, designing a tool that inserts cleanly into teachers’ workflows under the new reporting framework requires deeper operational insight into those workflows than I can currently claim. Second, I suspect actual AI use in report writing already runs well ahead of what surveys capture, which makes demand for a paid tool too uncertain to carry the cost of building one properly. Third, and most importantly: everything the two teachers told me points at problems that are systemic rather than individual. Comment boxes, review chains, records lost between schools: these are long-standing system design questions, not products of the current reform, and individual tooling cannot fix them. The right answer operates at a different level.
This edition also carries one correction from the March analysis: a widely reported statistic that 96 per cent of teachers rely on free AI tools proved, on direct verification against the NZCER source, to misattribute a different finding. The published figures are drawn from the report directly.
A second edition of this research, substantially expanded in scope, is in preparation.
— Greg Williams
31 July 2026
A reader survey now accompanies this article — for teachers, home educators, and parents — gathering structured evidence for the second edition. It is anonymous by default, runs on New Zealand infrastructure, and its methodology is published on the survey page itself. The document-request wording at the end of this article has also been aligned with how requests are actually handled: both volumes are sent automatically on request through the form.
Are you a teacher, home educator, or parent? The second edition needs your experience.
A short survey — about 3 to 8 minutes, anonymous by default, no names and no student information — is now gathering evidence for the second edition of this research, which will be provided to the Minister of Education. Whether you write reports, review them, home-educate, or read what comes home, your answers feed that evidence directly.
Take the survey →This is the web edition. The full paper (Volume I) and the supporting evidence base (Volume II — the V2-S references throughout) are available on request; see the end of this article.
Executive Summary
New Zealand is entering what is arguably the most significant reform of school reporting in over twenty years. From 2026, every student report for Years 0–10 must include five mandated elements: progress descriptors, a narrative explanation grounded in assessment data, assessment results from standardised tools, a visual representation of progress over time, and attendance records. This reform coincides with the rollout of SMART (a new standardised assessment tool replacing e-asTTle), new curriculum requirements, compulsory technology curriculum from 2027, and ongoing professional learning demands.
Teachers are absorbing these changes in a working context that is already under significant pressure. TALIS 2024 data indicates that New Zealand full-time teachers work 47.5 hours per week, the second-longest in the OECD behind Japan (all NZ TALIS 2024 figures carry a participation-rate caveat; see Section 2). Administrative work alone accounts for 4.1 hours per week against an OECD average of three. Sixty-five per cent of NZ teachers cite changing requirements from authorities as a source of stress, and forty-three per cent believe too many change initiatives are introduced at their school. Peer-reviewed research has established that marking and assessment writing is the single task most strongly associated with poor teacher wellbeing across English-speaking countries, including New Zealand.
Consumer AI is already present in schools: forty-six per cent of primary teachers in a nationally representative NZCER survey reported using AI tools, and a separate study of AI-engaged teachers found sixty-nine per cent using generative AI weekly. In either case, the use is largely ungoverned: most teachers rely on free tools, with roughly three-quarters of surveyed teachers reporting no school-funded access to premium versions; only 8% of schools have a policy for teachers’ AI use; and no Ministry framework exists for AI-assisted report narrative generation.[12]
The School Report Narrative Engine is a concept for a governance-first, calibration-driven tool that consolidates multiple evidence registers through a five-element architecture into individually differentiated, school-voiced, privacy-compliant report narratives. Its five structural elements are: teacher voice calibration, student context differentiation, whānau reception awareness, a longitudinal child development register, and peer review chain calibration. Professional judgement remains with the teacher as the interpretive authority: the engine drafts; the teacher confirms.
The value proposition is relationship transformation: transform the teacher–whānau relationship while spending less total time on communication than teachers currently spend on reports alone. An author-developed time model estimates the current reporting burden at approximately 42 hours per teacher per reporting cycle (84 hours annually). Even conservative projections suggest the engine could return approximately one full working week per year — time that returns to teaching, professional development, and family.
This document is designed to start a conversation with practising teachers. The architecture is built on assumptions that require practitioner validation. This analysis is offered as a foundation for confirmation, correction, and extension by those closest to the work.
1. The Reform Context
In February 2026, Education Minister Erica Stanford announced a nationally consistent reporting approach for Years 0–10 students, mandating five elements in every report: one of five progress descriptors — Emerging, Developing, Consolidating, Proficient, or Exceeding — for reading, writing, and mathematics; a narrative explanation of why that descriptor was chosen and how parents can support next steps; assessment results from standardised tools; a visual representation of progress over time; and an attendance record.
The framework was developed with principals’ associations and teachers, and trialled in 85 schools involving approximately 12,000 student assessment engagements.[20] It is well designed. The question this document addresses is not whether the reform is sound — it is whether the implementation environment can absorb it.
Alongside it, teachers are learning SMART, a new standardised assessment tool under a five-year, approximately AUD 21 million contract,[1] replacing e-asTTle for twice-yearly assessment of Years 3–10.
V1-Fig-1: The Convergence — five simultaneous demands converging on a single teacher. Source: Author analysis. Evidence references: TALIS 2024, MoE 2026. See V2-S3, V2-S4.
The reform does not land in isolation: teachers are simultaneously absorbing SMART, the new reporting framework, legacy data migration, a compulsory technology curriculum from 2027, and Government Target 72 (80% of Year 8 students at expected curriculum levels by December 2030). No single demand is unmanageable — the compound effect is the problem, and TALIS 2024 confirms New Zealand teachers already report change fatigue well above the OECD average.
2. The Working Context
The OECD’s Teaching and Learning International Survey (TALIS) 2024 provides the most comprehensive international dataset on teacher working conditions. The findings are striking, though they carry an important caveat: New Zealand did not meet TALIS 2024 participation rate standards, and all NZ-specific estimates should be interpreted with caution due to higher risk of non-response bias.[2]
With that caveat noted: New Zealand full-time teachers report working 47.5 hours a week — second-longest in the OECD behind Japan’s 55 — against an OECD average of 41.[3] The gap is not teaching time (21.5 hours against 22.7) but non-teaching work, especially administration (4.1 hours a week against an OECD average of three). More than 30% of teachers experience stress “a lot”, among the highest rates internationally and up five-plus points since 2018;[4] sixty-five per cent cite “changing requirements from authorities” as their top stressor — precisely what the 2026 reform represents.[3]
One finding deserves particular attention: peer-reviewed research by Jerrim and Sims (2021), analysing TALIS 2018 data across five English-speaking systems including New Zealand, found marking and assessment writing is the single task most strongly associated with poor teacher wellbeing[6] — not total hours, specifically the act of writing assessments and marking. Report writing is the concentrated form of that task, repeated 25 times a cycle.
International evidence reinforces this: a commissioned Scottish workload study found reporting processes extend working hours, and that reducing extensive after-hours marking may improve, not compromise, outcomes.[10]
One counterpoint: the NZCER 2024 National Survey found 90% of teachers enjoy their work and 75% report good morale — not every indicator is worsening; the profession cares deeply about the work but is structurally overloaded in specific, identifiable areas.
3. The Technology Context
AI has already arrived in schools. The NZCER 2024 National Survey — a nationally representative sample of 639 teachers from 148 schools — found 46% reported using AI tools in their teaching.[11] A separate, self-selecting NZCER sample of 266 primary teachers found 69% using generative AI weekly.[12] Both point the same way: AI use is widespread, growing, and largely ungoverned.
Most rely on free chatbots: 73% report no school-funded access to premium tools, and only 8% of schools have a policy governing teachers’ AI use.[12] The Ministry has published principles-based guidance on AI in assessment, but no framework exists for AI-assisted report narrative generation.
Current market offerings for AI-assisted report writing operate as prompt-wrapper comment generators: a teacher selects a student, enters a few descriptors, the tool drafts a comment. That addresses the speed of a first draft while leaving unaddressed what makes a report meaningful — the teacher’s voice, the child’s context, the family’s reception needs, institutional review standards, and the longitudinal continuity of the child’s story.
The gap is architectural, not incremental. No New Zealand product, at the time of this analysis, consolidates multiple evidence sources through multi-axis calibration under governance into individually differentiated narrative reporting aligned to the five-element framework.
The governance gap is acute given student report data involves minors. A hallucinated claim about a child’s progress is not an inconvenience — it is a professional liability event for the teacher and potentially harmful to the child; that risk is structural to any AI-assisted reporting approach, including this one. The trust architecture must be at least as rigorous as anything applied to adult-facing AI tools, because the human in the loop is time-poor and reviewing output in batches of 25.
4. The Time Burden
To understand the scale of the reporting task, the author developed a per-student time model based on stated assumptions — not measured data, but a structure for practitioner validation.[13] The model estimates the current reporting burden for a 25-student class at approximately 42 hours per cycle, including an IEP adjustment for the roughly 16% of students with additional learning needs. As a cross-check, two practising school leaders independently estimated approximately 50 hours per cycle — the same order of magnitude.[14]
The 42 hours break down across four phases: data collection and collation (16.8 hrs), report narrative writing (17.9 hrs), the peer review chain (5.1 hrs), and administrative completion (2.2 hrs) — 84 hours annually across two reporting rounds.
V1-Fig-5: Where the 42 Hours Goes — the per-cycle reporting workload by phase. Source: Author model, stated assumptions.
The model projects three scenarios for the engine’s impact once calibrated and operational: conservative (48.4 annual hours, 35.6 hours saved), moderate (33.0 hours, 51.0 saved), and optimistic by the third cycle (24.8 hours, 59.2 saved). Even the conservative scenario returns approximately one full working week a year — time for professional learning, differentiated planning, informal parent communication, or simply going home before dark during reporting season.
V1-Fig-6: The Time Dividend — projected scenarios against the 84-hour annual baseline. Source: Author projections; assumption-based, with low confidence on savings projections.
The institutional picture is larger still. For a 12-classroom primary school, the estimated total reporting burden is approximately 565–590 hours per cycle — over 1,100 hours a year across teacher direct time, syndicate leader review, deputy principal oversight, principal review, and coordination overhead.[15] These are illustrative estimates requiring practitioner validation, but the order of magnitude — roughly 140 eight-hour working days, more than half a full-time teaching position — is telling.
5. The Review Chain
The review chain is not merely an overhead cost. It serves a dual function that most workload analyses overlook — a quality-assurance track (syndicate leader, deputy principal, principal review) and a professional-mentoring track through which institutional voice is shaped and individual teachers develop their reporting craft.
The full analysis explores this in depth — request the complete paper.
6. Beyond Faster Reports
The value proposition of the School Report Narrative Engine is relationship transformation, not productivity.
Under the current model, the formal reporting cycle is the primary channel through which schools communicate a child’s progress to their whānau — two touchpoints a year; a parent finds out in July what happened in February. The engine enables a dual-output model: every real-time classroom observation captured through its mobile input channel can feed both the formal reporting log and an optional informal note sent home the same week, reviewed and sent by the teacher. (The mobile capture pipeline carries its own open privacy question, recorded in Volume II, Section 10.) This changes the whānau relationship from transactional to relational — a parent who has received three or four informal notes across a term arrives at the formal report already trusting the teacher’s understanding of their child.
The Equity Dimension
The transformation potential is not distributed equally, and that is precisely the point. Teachers in lower-decile schools — working with Pasifika families, recent migrants, or families facing poverty and housing instability — have the hardest time building relationships and the most to gain from a tool that makes warm, culturally attuned parent communication achievable within the working day. The greatest transformation potential sits where the relationship gap is widest: this is an equity argument, not a productivity one.
Making the New Way Achievable
The 2026 reform asks teachers to do more with new tools in a context where they are already under pressure. Every hour returned from reporting is an hour available for learning SMART, understanding progress descriptors, or professional learning. The engine makes the new way achievable, not the old way faster.
A Note on Scope
This document analyses one component of teacher workload: the reporting burden. It does not claim that reporting efficiency addresses the structural conditions of teaching — class sizes, staffing ratios, pay, or the demands collective agreement negotiations rightly address; tool-based efficiency and structural workload reduction are complementary, not substitutive. Nor does it resolve who should fund such tooling — a model where teachers pay personally for professional infrastructure is itself a workload-equity question collective processes rightly examine.
7. The Architecture: Five Structural Elements
The engine is built on five structural elements, each requiring distinct calibration, addressing dimensions current market offerings leave untouched.
V1-Fig-8: The Five-Element Architecture — an at-a-glance overview of the five structural elements detailed in this section.
Element 1: Teacher Voice Calibration
Each teacher writes differently, and reports carry professional identity. The engine ingests a teacher’s previous reports as calibration data and learns their voice — sentence structure, vocabulary range, how they balance specificity with warmth — so output reads like the teacher who knows the child.
Element 2: Student Context Differentiation
Every child requires structurally different narrative, not just different data. A child on an IEP requires a fundamentally different narrative structure than a child who is exceeding expectations; the Ministry explicitly states that children with additional learning needs should not be automatically assessed as Emerging.[21]
Element 3: Whānau Reception Awareness
How you communicate progress to a Pasifika family differs from how you communicate to a recent migrant family. Cultural context is central to Te Tiriti o Waitangi obligations and the New Zealand Curriculum, and the teacher holds this knowledge through the whānau relationship — no consumer AI tool even knows to ask about it.
Element 4: Longitudinal Child Development Register
At every school transition, a teacher’s accumulated understanding of a child either does not transfer or arrives as a thin data extract. The engine maintains a child-owned longitudinal register, keyed to the National Student Number (NSN), that persists across transitions and bridges assessment-framework discontinuities (Te Whāriki Learning Stories in ECE through to NZC progress descriptors). This element activates only once a dedicated trust architecture is resolved.
Element 5: Peer Review Chain Calibration
The engine profiles each reviewer’s correction patterns and pre-applies institutional standards at generation time (reviewer awareness and consent for this profiling is an open design question; Volume II, Section 10), shifting the reviewer’s time from catching predictable compliance gaps to genuinely developmental feedback. The core performance metric is the first-pass approval rate through the review chain.
V1-Fig-7: Five Elements — Gap Analysis. Source: Author assessment of the current market category at the time of this analysis.
The contrast is stark: today’s market tools offer generic tone selection only, leaving whānau reception, longitudinal continuity, and review-chain calibration unaddressed — the five gaps the concept architecture is built to close.
Privacy by Non-Ingestion
V1-Fig-9: the privacy architecture, expressed as its own governing principle: non-ingestion rather than access control. Open privacy questions are recorded in Volume II, Section 10.
The child’s name never enters the engine. The National Student Number (NSN) is the primary key for all student context; the teacher knows which NSN belongs to which child, the engine does not need to. The reporting log lives on the teacher’s own school drive — no hosted platform, no data sovereignty question. In a context where SMART was, at the time of this analysis, under active OIA scrutiny on data sovereignty questions,[19] this architecture is designed to minimise re-identification risk rather than to manage it after the fact. That comparison is scoped to offshore-transmission and hosted-platform risk; whether NSN-keyed context itself constitutes personal information under the Privacy Act remains an open question (Volume II, Section 10).
8. How It Works, Who Buys It
The buyer is the individual teacher, not the school or the Ministry. Schools will not move without MoE approval, and MoE moves at its own pace — but a teacher facing 30 reports, a two-week deadline, and the review chain’s iteration cost will pay for something that gives them their evenings back.
Teachers who use AI already pay for their own subscriptions; the product is the calibration workspace and profiling tools that make an existing AI subscription produce professional-grade reports instead of generic comments, priced so a teacher does not need institutional approval to buy it.
The operating cycle: the teacher maintains a structured reporting log; each round, they share it with the engine, which reads it cold, drafts reports grounded in the accumulated context, and returns an updated log for the teacher to review, edit, and submit. The engine is stateless between sessions — the log is the continuity mechanism.
Institutional adoption follows from the ground up: when enough teachers in a staffroom are using it and the principal notices reports arriving consistently cleaner through the review chain, adoption spreads organically — no procurement tender, no Ministry sign-off at market entry.
This ground-up path sits in deliberate tension with the governance case made earlier, and the tension deserves naming rather than glossing: a teacher adopting any AI reporting tool individually — including this one — should disclose that use to school leadership, and the concept treats that disclosure as a condition of professional use.
This is not a hosted platform: the product is methodology and tooling — calibration workspace, voice profiling, context framework, whānau guidance, and review chain profiling tools, built on the school’s own drive and management system.
This document is built on assumptions — some grounded in published research, others professional estimates — offered for testing against the experience of people who actually write school reports, review them, and live with the consequences. This document is designed to start that conversation.
References
1. GETS tender MOE29225. Janison Solution Pty Ltd, five-year contract, approximately AUD 21 million. Awarded December 2025.
2. All NZ figures from TALIS 2024 carry a participation rate caveat. NZ did not meet TALIS response rate standards; estimates should be interpreted with caution due to higher risk of non-response bias. OECD, Results from TALIS 2024: Country Notes — New Zealand (2025).
3. OECD, Results from TALIS 2024: Country Notes — New Zealand (2025). Full-time lower secondary teachers. Japan reported 55 hours.
4. OECD, Results from TALIS 2024: Chapter 2 — Thriving in Teaching (2025).
6. Jerrim, J. & Sims, S. (2021). ‘When is high workload bad for teacher wellbeing? Accounting for the non-linear contribution of specific teaching tasks.’ Teaching and Teacher Education, 105, 103432. Peer-reviewed. Based on TALIS 2018 data across five English-speaking systems including New Zealand.
10. Hulme, M. et al. (2024). Teacher Workload Research Report. University of the West of Scotland / Educational Institute of Scotland. ISBN 978-1-903978-77-1.
11. Li, M. et al. (2025). Primary school teachers’ perspectives from the 2024 National Survey. NZCER. DOI: 10.18296/rep.0081. Nationally representative: 639 teachers from 148 schools.
12. Coblenz, D., Dong, J. & Gibbs, B. (2025). Generative artificial intelligence in Aotearoa New Zealand primary schools. NZCER. Voluntary sample of 266 primary teachers. The authors acknowledge the sample was weighted towards teachers with an existing interest in AI.
13. Author’s model based on stated assumptions. Per-student breakdown in Volume II, Section 6. Designed for practitioner validation, not presented as measured data.
14. Lovell, O. (2021). ‘A better way to do school reports?’ Practitioner survey; two primary school leaders independently estimated approximately 50 hours per teacher per reporting cycle.
15. Author’s illustrative model. Full assumptions in Volume II, Section 7. The figures represent estimated institutional overhead above individual teacher direct time.
19. The Treasury (2026). Official Information Act response 20250889, 10 February 2026 — Treasury advice to the Minister of Finance on the SMART project. treasury.govt.nz.
20. Office of the Minister of Education (2026). ‘Assessment and reporting changes for parents shine light on learning.’ Media release, 2 February 2026. beehive.govt.nz.
21. Ministry of Education. Assessment and reporting guidance for 2026. Tāhūrangi, tahurangi.education.govt.nz.
This is the published concept brief. The full paper goes deeper.
Volume I is the complete narrative, with full references. Volume II is the supporting evidence base — data tables, sensitivity analyses, bibliography, and open privacy questions. Both are sent by email, automatically, on request through the form below.
Request the full paper →Edition. First Edition — 12 March 2026. Prepared March 2026; published 30 July 2026 under Steko’s structured research-publication method. The dated editorial note above records developments between preparation and publication.
How this article was produced. This article was produced under Steko’s structured research-publication method, a governed production method for research articles. The method requires argument definition, source inventory with gap analysis, structured section architecture, draft production, source-faithful resonance checking, and a multi-pass adversarial review (hostile-reader archetypes, substance traceability, register and brand compliance, and AI-fingerprint diagnostics) before publication.
What the practitioner brought. The research question and its origin in years of first-hand accounts of reporting-season workload; direction of every search; evaluation of every source; every positioning decision; approval of every claim; the two-teacher validation round; and publication judgement throughout.
What the production engine brought. Structured research synthesis; per-section drafting to architectural specification; construction of the time model to stated assumptions; and the adversarial-review instrumentation across all four passes.
Powered by Claude Opus 4.6, for the research and analytical production (March 2026), and Claude Fable 5 with Claude Sonnet 5, for the publication and editorial production.
Hostile-reader review. Six archetypes tested at publication — a Ministry policy analyst, a union research officer, a school principal, an EdTech market participant, an academic reviewer, and a privacy advocate. Every objection class carries an in-document defence.
Substance traceability. Every load-bearing factual claim carries a named source. The author spot-checked the highest-stakes claims against the primary sources at publication, including a correction to one widely reported secondary statistic, disclosed in the editorial note above.
Reader feedback and correction. If you identify a factual error or a claim that does not hold, write to thinking@steko.co.nz. Valid corrections are made promptly and acknowledged.
This article was produced with substantial AI assistance under Steko Consulting’s structured research-publication method, covering research synthesis, draft generation, and structural composition, as set out in the colophon above. All material has been reviewed by the named author, who takes full professional responsibility for its contents.
© 2026 Steko Consulting Limited. All rights reserved. This article and the underlying analytical compilation (Volume I and Volume II) remain the intellectual property of Steko Consulting Limited. Reproduction or redistribution beyond fair citation requires written permission; requests for the full paper are handled via the request form as described above.