Grading the machine's lesson plan: building developmental judgment by critiquing AI activities

Child Development and Educational Studies English Beginner 75 min Free. The published study used a free tier deliberately so every student had the same tool.

The situation

My students can recite the developmental guidelines on an exam and then hand me a practicum plan asking three-year-olds to sort by two attributes and write their names on a chart. The gap is between knowing the principle and seeing the violation in front of them. So I stopped asking them to write a plan first.

Steps

  1. Generate plans with a prompt that does NOT say developmentally appropriate

    Any free tier

    Specify age, setting, a learning goal and a time block, but deliberately omit any developmental-appropriateness language. Save the raw output verbatim before class.

    What you only learn by doing it: This is the whole trick. Put “developmentally appropriate” in the prompt and the model self-corrects into something bland but defensible, and your students have nothing to find. Leave it out and the flaws arrive on their own. Generate five or six and keep the three with the richest problems.

  2. Have students work in pairs against a named source, not from memory

    NAEYC guidelines, your observation rubric

    Pairs annotate the plan, and every objection must cite a specific guideline or developmental milestone by name. The talk between partners is where the reasoning surfaces.

    What you only learn by doing it: Require a minimum number of specific objections — three is about right. In the published study, 16 of 50 assignments recorded no disagreement at all with the AI's suggestions. Fluent, confident prose reads as authoritative to novices.

  3. Forbid regenerating or follow-up prompting

    None — this is a rule

    Students work with the first output, unedited. The point is to make them do the pedagogical work rather than prompt their way to a better plan.

    What you only learn by doing it: Students will ask if they can “just tell it the kids are three.” Say no, and explain why: in the field they will be handed curriculum written by someone who has never met their children, and the job is to adapt it, not to request a better version.

  4. Revise the plan and log every change with a reason

    Google Docs with tracked changes

    Pairs rewrite the plan and keep a change log: what they changed, which principle drove it, what they left alone.

    What you only learn by doing it: Push hard on deletion. In the study, content was removed in only 10 of 50 assignments while activities were added in 33. Novice teachers expand plans because adding feels like effort — but over-scheduling a three-year-old's morning is the most common real-world failure. Make “what did you cut, and why” a required line.

  5. Debrief on the pattern, not the individual plans

    Whiteboard or shared doc

    Collect categories of flaw across the class: developmental mismatch, materials feasibility, time and ratio assumptions, engagement mismatch. Naming the categories turns a one-off critique into a reusable screening habit.

    What you only learn by doing it: Developmental inappropriateness will dominate — it appeared in 21 of 50 reflections. Have a second lens ready for what students under-detect, especially adult-child ratio and transition time, which AI plans get wrong constantly and novices almost never flag.

Where this breaks down

The study's own authors caution that the critical engagement students display may be partly an artifact of being asked about disagreement, and that reflections stayed abstract — no revised plan was ever taught to actual children, so nothing verifies the revisions were better.

There is a real risk of teaching cynicism instead of judgment. Students can learn to produce ritual objections without developing the underlying knowledge, which is why every objection must be anchored to a named source.

Be careful the exercise does not implicitly endorse using AI to write practicum plans. Students will notice you generated one in eleven seconds; say plainly what you think about that.

There is an equity dimension: students with prior early-childhood work experience find flaws instantly while students entering cold may see nothing wrong. Pair deliberately rather than letting pairs self-select.

Provenance: documented practice. Özüdoğru & Delen (2026), “Beyond Acceptance: How Early Childhood Preservice Teachers Critically Engage with GenAI Feedback in Lesson Planning,” Education Sciences — 50 preservice teachers, 25 pairs, two assignments four weeks apart. The source study used a national curriculum template rather than NAEYC; that framing is our adaptation.