The AI runs the test, the student defends the assumptions
The situation
In intro stats the computation stopped being the hard part years ago. A student can paste a dataset into a chatbot and get a t-test, a p-value and a paragraph of confident prose in forty seconds. What it will not reliably do is notice the design was paired and it ran an independent-samples test. So the assignment inverts: the AI must compute, and the grade is entirely on the judgment.
Steps
-
Hand out a dataset with a deliberate design trap
A CSV from data.gov, CDC, or your own institution
Choose data where the naive test is the wrong test: repeated measures that look like two independent groups, strong skew with small n, a lurking confounder, clustered observations. Give the research question in plain English with no guidance on procedure.
What you only learn by doing it: Put the trap in the data structure, not the arithmetic. The reliable one is paired data presented in two columns with no ID variable — the model will almost always run an independent-samples test, because that is what the shape suggests and it has no access to how the data were collected.
-
Require the AI to compute, and the student to keep the transcript
Google Colab (Gemini features and Data Science Agent) Julius AI
Students prompt the model to perform the analysis. The full transcript — every prompt and every response — is a required attachment. This makes the process visible and removes any incentive to hide the tool.
What you only learn by doing it: Require the prompts, not just outputs, and read them. The prompt is where you see whether the student told the model the design. Students who supply design context get a correct analysis; students who do not get a confident wrong one, and showing the two transcripts side by side is the most efficient teaching moment in the unit.
-
Run the assumption audit as a separate graded document
For each assumption the procedure requires: state it, state how it was checked, show the evidence, give a verdict with consequences. Crucially, also evaluate whether the procedure the AI chose matches the study design at all.
What you only learn by doing it: Add a mandatory line the model almost never volunteers: how were these data collected, and does that match what this test assumes? Everything else it will happily check when asked. The design-to-procedure match is the one thing it structurally cannot verify from a file.
-
Make the students force an error and document it
The same assistant
Students deliberately re-prompt with the design information omitted or misstated, capture the wrong analysis, and write up how it differed and whether they could have caught it from the output alone.
What you only learn by doing it: Do this after the correct analysis, never before — students who see the failure first conclude the tool is useless and disengage. What makes it land is that the wrong output looks exactly as authoritative as the right one. Have them note what would have tipped them off; usually the honest answer is nothing.
-
Write the interpretation in context, with a limitation the AI did not raise
A word processor, and no AI permitted on this section
A short interpretation for a non-statistical decision-maker: what the result means, what it does not license, and at least one limitation the model did not mention. This section carries the most weight.
What you only learn by doing it: Specify that generic limitations do not count. Models produce competent generic caveats about sample size and generalisability; require something concrete about this dataset or this sampling frame, or you get thirty near-identical paragraphs about correlation and causation.
Where this breaks down
The model selects a procedure from the shape of the file rather than the design of the study, then writes an interpretation whose confidence is unrelated to whether the procedure was appropriate. Published work documents students accepting well-formatted output without interrogating procedure choice.
Taught carelessly this produces students who can operate a tool and narrate its output with no model of inference underneath. Pair it with some by-hand or small-n work where they compute something and see why the formula has the shape it does.
The audit and the interpretation are both vulnerable to being AI-written, which is why the required transcript, the supplied audit template and the in-class interpretation matter more than an honour code here.
Outputs are not reproducible between students or sessions, so build the rubric around the reasoning rather than a specific expected number.
Provenance: documented and adapted. Grounded in Schwarz, Teaching Statistics 2025, and “Students' Statistical Thinking When Using Generative AI,” JSDSE 2025, both documenting the shift of assessable work toward interpretation and critique. The forced-error step is adapted from Kim & Kim, ZDM 2024, which found students confronting contradictory outputs became markedly more critical.