Kali Trzesniewski · Research program

The research program · 03

Build the environments, then scale what works

Effective use has to be learned somewhere, and measured somewhere. Both mean building the settings first, which runs into what instructors can carry, and whether they see the evidence about their classrooms.

Why the environments come first

The prerequisite, twice over. People do not become effective users of Generative AI (GenAI) by being told the rules. It takes tasks worth doing, goals worth reaching, room to decide when the tool serves the goal and when it does not, and the chance to notice the difference — which is to say it takes a learning environment built for it. That is the developmental case, and it is the ordinary one: skills develop where they are practiced.

The methodological case is the one the field keeps missing. A measure of effective use cannot be validated in a setting that never elicits the full range of use: where every assignment has one right answer and the policy is enforcement rather than instruction, offloading is the rational choice, avoidance is the safe one, and thinking with AI has no reason to show up. The environments are not where the science gets applied afterward. They are the condition under which the measures need to be developed.

So why are they rare? Not because instructors do not want them. Two constraints, replicated across very different settings.

The first is labor. Building that kind of environment takes preparation before the course and recurring work at the scale of a class, every week, all term — and instructors are not given that time. Programs that assume the problem is motivation keep offering evidence and workshops to people who are already convinced.

The second is the feedback loop. An instructor rebuilding a course needs to know what is actually happening in it, and the evidence about their classroom either never reaches them or arrives in a form they cannot act on. Without that, redesign is guesswork and improvement does not compound.

Constraint 1: the labor

The same lesson, across very different contexts. A 2,097-school randomized trial in Indonesia and a National Science Foundation (NSF)-funded national instructor fellowship in the U.S. (award #2201928) taught the same lesson in settings that differ in almost every way: cultural norms around direction versus autonomy, middle school versus college, how much freedom teachers have to change their own classrooms. In Indonesia, the version of the program built to change daily classroom practice was the version schools could not implement: training was do-it-yourself at national scale, and schools complied at roughly half the rate of the schools given only the curriculum, naming burden and time, while more than two-thirds of teachers recommended the program. The environment still moved: teachers' beliefs about failure shifted, and students — disadvantaged students most of all — reported classrooms more focused on learning progress and effort. In the U.S. fellowship, instructors who volunteered (motivation removed as the explanation) adopt the low-effort verbal practices nearly universally and stop short of the high-effort structural changes that alter what students experience. Opposite selection, opposite cultures, same constraint: implementation capacity, not conviction. Instructors need time, tools, and support. A program that does not budget for the teacher's labor scales its own failure.

How I am using GenAI to help: the workload tests. I helped two colleagues rebuild their undergraduate courses around authentic assignments and assessments, so that students were motivated, engaged, and saw the value in what they were learning. The courses were designed to make the purpose of the learning clear and to connect it to learning students value — designed around valuing the process of learning, and valuing every voice in the room as contributing meaningfully to the learning environment. The workload test: whether GenAI could carry enough of the recurring labor that a redesign like that survives the whole term. The instructors kept content authority, reviewed everything students saw, and stayed the voice of their courses. Here is what GenAI carried, successes and struggles included — I am still learning.

A complete redesign of the course structure: the first step was to translate the course learning goals into authentic assignments that provide opportunities for students to learn, make mistakes, and improve. I had Claude create a specialized knowledge base on authentic assignments and assessments, growth mindset, belonging, and purpose of learning, as well as the science of learning and motivation. Then I brainstormed ideas, and the instructors chose their focus and made edits. Claude then broke the assignments down into weekly tasks and templates for the instructors and students, suggested grading weights, proposed timelines, and built out instructor and student packets. The instructors reviewed everything and made plenty of changes — but they got to start with editing instead of a first draft. Unfortunately, I did not have permissions to connect Claude to Canvas, our campus's course platform, so the instructors still had to do the full Canvas building.

The opportunity to hear from every student, every week, and adjust teaching in response: GenAI carried the labor of creating first drafts of weekly exit tickets (the short reflections students write each week), summarizing the themes each week, answering questions the instructor had about what was written, surfacing examples when more information was desired, providing first drafts of messages to send to the course, examining patterns across weeks, and suggesting adjustments to teaching or to the resources provided to fill gaps.

What the students said: across the courses, a large majority of the students surveyed recommended keeping the redesign, with first-generation students rating it higher than their continuing-generation peers.

Constraint 2: the feedback loop

The gap. Institutions collect data on student experience continually, and almost none of it changes what happens in a classroom. Reports sit unread. Data remain unanalyzed. Results come back as abstracted aggregates, disconnected from the context that produced them and from the practices that could change them. The richest material — what students say in their own words — is either never collected, reduced to a couple of cherry-picked quotes, or ignored. For an instructor who wants to teach better, the process fails at every step and never connects to a practice they could try on Monday.

Three attempts, and a decision to stop. The lesson arrived while delivering results from a course-experience measure — developed and still being validated under NSF award #2201928 — in a setting where faculty had not opted in. We did what the field does: results by email, a faculty-meeting presentation, a meeting with the committee responsible for teaching. Each attempt either left faculty overwhelmed by the volume or drew accusations of cherry-picking. The data were sound; nobody was reading them, and the conversations the reports did start were about questions the reports were not built to answer. We went looking for another way to communicate.

Every Voice, Every Classroom (Trzesniewski & Ebeler, 2026). The questions named what the reports got wrong: faculty wanted more, and wanted different, than what we picked out to tell them — they each wanted to ask their own questions and discover what is more interesting and useful for their own teaching and their own classes. So the design inverted: give them everything, and let them ask. Every Voice, Every Classroom is an interactive resource where instructors explore what their students actually said — an engaging entry page with student voices leading, and behind it the full curated, redacted set of student responses and ratings, connected to a library of evidence-based practices, queried in plain language. Every answer comes back in the same shape: student voice first, then the theme, the numbers, students' own advice, and the practices that match the question. The goal is the same one that runs through this page — evidence returned in a form that motivates people to want to act on it.

It reached faculty in a way the written reports had not. One who had been among the hardest to engage described being stopped short by students' words about assessment, reflecting on how their own practices might produce that experience — and went on to propose using the resource for professional development, departmental discussions on teaching, and curriculum planning.

Where this goes

The direction is a feedback loop with a memory: course-experience surveys that become conversational, where early randomized comparisons suggest AI-led interviews elicit more detailed and informative responses than standardized administration (Xiao, Zhou, Liao & Zhou, 2020; Barari et al., 2025, preprint; Wuttke et al., 2026, preprint), and reports that become dynamic: generated on demand, tailored to an instructor's own course and questions, able to compare across courses and remember what was tried last year, what worked, and what didn't. The pieces belong together rather than beside each other, and building them together — kept under evaluation as they scale — is the program's third line of work.

References