CS 124 is introductory computer science (CS1) at the University of Illinois, enrolling over 2,000 students per year. 80% are non-majors. Here’s what we’re doing and why.
Geoff Challen has been teaching the course since Fall 2017, and along the way we’ve learned a lot about what works when teaching intro CS at scale. A recent talk at the University of Sydney overviews many of the course innovations described below.
This is where things have gotten really interesting.
The journey. In Fall 2024, we began allowing students to use AI on the course project as an experiment. Over the summer of 2025, Geoff tested whether Claude could complete the Fall 2025 project from test suites alone—and it could, completing almost everything with minimal human intervention.
In Fall 2025, we tried a compromise: keeping the same project structure while allowing AI. But it didn’t work. Traditional programming assignments hinge on a specification that students translate to code, and AI coding agents have become too good at that translation step. The tension between precise specifications—needed for fair automated grading—and imprecise specifications—needed to prevent AI from doing all the work—proved irreconcilable.
The key insight. If AI handles the translation from specification to code, the valuable human contribution is formulating the idea and specification itself. That realization drove the redesign.
Starting in Spring 2026, every student designs and builds something of their own, working with AI coding agents—specifically Claude Code—throughout the process. Students aren’t translating our specification into code—they’re creating their own specification and learning to communicate it effectively to an AI agent.
The preparatory pipeline. We don’t just hand students an AI tool and wish them luck. The project unfolds as a carefully scaffolded sequence of weekly discussion-section activities that carry students from “I have no idea what to build” to a shipped app:
The workshops rotate among three formats: partner pairing, pair critique, and a status update followed by pairing up to share progress and decide what to do next. Students can build a web, Android, iOS, or native desktop app, and every project must be usable by another student, either through a link or by handing over a phone or a supervised laptop. Each activity builds on the previous one, moving students from “I have no idea what to build” to an app their classmates can use.
When every student builds a unique app, traditional autograding is impossible—there are no shared test suites to run. Manual rubric-based grading is infeasible at our scale and inevitably subjective for creative projects. Instead, we grade the project on measured effort from AI coding agent session activity.
The project repository carries Claude Code hooks that stream each session’s events to the course as they happen. Each prompt a student sends earns the time Claude spends working on it, the time it takes to read the reply and write the next prompt, and two minutes to think. Time away from the conversation does not count, and a permission dialog left open earns at most two minutes. The rule is calibrated against real sessions, and no language model is involved in the measurement.
Students earn full project credit by accumulating 24 measured hours between October 6 and December 10, a pace of about 3 hours per week, with a cap of 8 credited hours per week. Working with Claude Code during a Tuesday section counts like any other time; attendance by itself earns nothing. The weekly cap discourages cramming and encourages sustained engagement. Students see their hours update as they work, and can review every session and the minutes each turn earned. The grading criteria are completely transparent—invest the time, earn the credit.
Learning objectives for agentic development:
Agentic development doesn’t replace classical programming—it complements it. We weight proctored quizzes at 60% of the grade, taken without AI in a computer-based testing facility. Across 14 weekly quizzes, students complete 36 proctored programming challenges—far more proctored coding than most CS1 courses require. The project is 20% of the grade and is where students practice agentic development. This split lets us teach both classical programming and agentic development without one undermining the other. Computational thinking unites the two—whether students are writing code themselves or directing an agent, they need to decompose problems, think algorithmically, and reason about correctness.
For the full story, see Geoff’s presentation to CS 124 students at the end of the Fall 2025 semester, and a follow-up talk on using and teaching coding agents.
CS 124 doesn’t hold lectures. Instead, students work through daily interactive lessons that mix text, runnable code examples, interactive walkthroughs—more on those below—short videos, and practice problems. Students engage at their own pace, which matters a lot when you have a wide range of prior experience in the room—from students who’ve never written a line of code to those with years of experience.
Daily engagement also achieves spaced repetition naturally. Rather than cramming before a midterm, students return to new material every day. Because students move at their own pace, they can slow down on topics they find confusing and spend more time with the material that challenges them. We’ve also built up explanations from multiple instructors, so if one explanation doesn’t click, there’s usually another that might.
A course without lectures has an obvious failure mode: students who skip the lessons. So the lessons’ interactive components are instrumented. Each video and walkthrough is divided into 8-second intervals, and an interval counts once it has actually played: rewatching adds nothing, and skipping ahead counts for nothing. Playback pauses when the player scrolls out of view, so a lesson left running in a background tab does not count either. An item counts as complete at 90%.
That allows us to give every student a report on how much of the lesson content they have reviewed, lesson by lesson, with one bar segment per video or walkthrough, colored from red (unwatched) to green (watched). Here is a sample, with made-up numbers modeled on a real student’s report:
You've reviewed only 17% of the lesson content.
Lesson completion is not graded, but it's difficult to learn without reviewing the material.
We don’t grade completion. The data isn’t strong enough for that, although Craig Zilles and Matt West’s work on Bayesian grading suggests ways it could become part of a grade. But we put the report at the top of the grades page, directly below the student’s running total and above every graded section, so students scroll past it every time they check their progress.
Plenty of students still don’t do “the reading.” But a student who is struggling now has to confront the most likely reason why, and knows that we can see it too. That has simplified many of our conversations with struggling students: almost always, the problem with their approach to the course is right there in the report, and the students whose difficulties are not explained by it are the ones most deserving of our time.
CS 124 has no final exam, no midterm, and no high-stakes assessments. Instead, students complete a daily homework problem and take a weekly 50-minute proctored quiz in a dedicated computer-based testing facility. Each quiz is worth only about 5% of the grade. Everything is autograded with immediate feedback, and students get unlimited attempts on programming problems—until the deadline or quiz timer runs out.
We’ve found that frequent small assessment may be the single most important component of student success. Students study regularly, we catch struggles early and reach out when someone does poorly on a quiz, and the low stakes reduce anxiety. Frequent data points also enable learning-focused policies that would be impossible with high-stakes exams: dropping lowest scores, allowing quiz retakes where students return to questions from previous weeks, and catch-up grading where doing better on a later quiz raises earlier scores.
Counterintuitively, this is also more rigorous. In Fall 2017, students wrote code in a proctored setting exactly once—on a paper final exam. Now they complete multiple autograded programming challenges every week, totaling 15 hours of proctored assessment per semester versus 3–4 previously.
Since Fall 2026, each quiz also comes back with feedback. Students have asked for years to review their quizzes, and because questions are reused, they never leave the testing facility. Instead, once a quiz closes, a student’s grades page points them at the lesson sections worth another look, drawn from how their programming attempts went and which multiple-choice items they have not yet answered correctly. How that works is described under LLMs below.
For more on this approach, see Geoff’s talk on frequent assessment.
One of our more distinctive innovations is what we call interactive walkthroughs—recorded code editing sessions that students can replay and interact with. These aren’t videos. They’re actual editor replays: students see code being written character by character, can pause and edit the code themselves, and resume the walkthrough. It’s closer to watching someone code live, except you can rewind, speed up, and experiment along the way. You can see examples at learncs.online.
A student interacting with a walkthrough
An instructor recording a walkthrough
Frequent assessment creates a demand for problems. We needed a way to author them fast and accurately.
The insight behind our autograder, Questioner, is that when autograding, the solution is known. This is fundamentally different from software testing, where only the desired behavior is known. So instead of maintaining three sources of truth—a description, a solution, and tests—the author provides just a description and a reference solution, and Questioner generates and validates the testing strategy automatically using source code mutation.
The old process took hours per problem with unknown accuracy. The new process produces several problems per hour with validated accuracy. We’ve authored over 700 problems since Fall 2020 across multiple question types—code writing, debugging, tracing, and conceptual questions. Questioner also evaluates code quality—not just “does it work?” but “is it good?”—giving students instant feedback on complexity, style, and efficiency.
The same mutation engine produces debugging challenges—erroneous examples, in the research literature—at a scale no instructor could write by hand. Students in CS 124 have submitted correct solutions to hundreds of problems. We collect those, clean them up, and confirm each still passes the problem’s tests. Then Questioner’s source mutation engine introduces one small bug into each: flipping a comparison, nudging a loop boundary, changing a literal, removing a statement. The result is a pool of over three million buggy programs across more than 440 problems in Java and Kotlin, each one a real student’s working code with one thing wrong.
Here is one, from a problem asking students to reverse a string:
The student’s original loop ran while i >= 0; the mutation flipped it, so the loop never runs.
The student’s job is to find and fix the bug without rewriting the code. They may change only as many lines as the mutation did—one line for over 90% of challenges, measured by a line-by-line edit distance—and the fix must pass the problem’s tests and lint checks. That constraint is the point: it forces students to read and understand someone else’s code, written in a style that is not their own, rather than replacing it with their own solution. It also exposes them to approaches they would not have taken.
On 20 days this semester, the daily homework includes a debugging challenge. Each challenge gives every student a different random starting point in the pool, and five fixes earn full credit for the day. Students get two skips per problem, for a program that they cannot make sense of. They also get two checks: a student who believes a challenge cannot be fixed under the constraints can ask us to re-run the original, unmutated submission against the current tests. If it no longer passes—because the tests have grown stricter since that student submitted it, for example—the instance is withdrawn for everyone and the student gets a new one without spending the check. Otherwise they are told the challenge can be solved, and keep working.
Our tutoring model is built around peer tutors—recent CS 124 graduates who took the course themselves. They remember what was hard, they know the material, and they’re motivated to help.
The core of the system is an online-first tutoring platform that provides immediate 1-on-1 support throughout the day. When a student has a question—at 9 AM or 9 PM—a tutor is usually available within minutes. No waiting for office hours, no standing in line.
Student-tutor online tutoring interaction flow
Staff are organized into tiers: volunteer assistants gaining experience, paid associate mentors, TAs, and head TAs. When a student struggles on a quiz, tutors proactively reach out to offer support. We’re not waiting for students to come to us.
Beyond the project, LLMs are woven throughout CS 124 to solve specific educational problems that arise when teaching 2,000+ students.
AI Teaching Assistant. Students need help outside office hours, and even with a large staff, 2,000+ students can’t all get 1-on-1 time whenever they need it. A chat assistant answers questions using course content while maintaining academic integrity guardrails—it won’t solve homework problems. It uses retrieval-augmented generation (RAG) to search transcribed lectures and walkthroughs, so answers stay grounded in what’s actually been taught rather than generic internet knowledge.
Automatic Transcription. Hundreds of hours of walkthrough recordings and videos need to be searchable and accessible. WhisperX runs locally to transcribe all audio with word-level timestamps, enabling both the search pipeline and student-facing transcripts and captions. This makes the entire content library accessible to students who prefer reading, need captions, or want to search for a specific topic across all recordings.
Semantic Search. With a large and growing content library, students and the AI assistant need to find relevant material quickly. Transcripts are chunked, embedded, and stored in a vector database for hybrid semantic and keyword search across all course content. A student asking “how do I use recursion with linked lists?” gets pointed to the right walkthrough segment, not just a keyword match. The same index is searchable directly from the site.
Course Content Over MCP. Students are already required to use a coding agent for the project, and were reaching for their own AI whether or not we pointed it at anything useful. Rather than compete with that, CS 124 publishes its content to the agent the student is already running over the Model Context Protocol. Each student connects with a single command and a personal key; their AI can then search lessons and walkthrough transcripts, read a full lesson, ask which lessons the next quiz draws on, and browse the practice problems by lesson or concept.
The design constraints are worth stating, because they are what make this defensible. An MCP tool description is advisory—a client model can ignore it—so nothing may depend on the student’s AI behaving. Every tool is read-only, exposes no grades or submissions, and is filtered by the same rules the website enforces: the student’s language, the release date of each lesson, and, for a homework problem’s solution walkthrough, its deadline or the student having solved it. No quiz questions are served at all. An agent holding the lessons writes better practice questions than the item bank would provide anyway, and questions a student has not already seen are better practice.
We think this is the more honest response to students using AI: the alternative is not that they use less of it, but that they use it against material we did not choose.
Quiz Feedback. Students have asked for years to review their quizzes, which we cannot offer because questions are reused and never leave the testing facility. Since Fall 2026, once a quiz closes, a student’s grades page shows a few concept-level pointers back into the lesson sections worth another look. Two sources feed it. For programming questions, a model receives the student’s attempts, the autograder output they saw for each, the reference solution, and a fixed menu of the lesson sections the quiz draws on, along with a short per-quiz note from the instructor about what students have covered so far. It is asked for up to four concept-level pointers, or none, each tied to a section from the menu. For multiple-choice questions, no model is involved: each item the student has not yet answered correctly maps to the section that teaches it, and the pointer clears once a retake scores the item. The map underneath both—which lesson section teaches which quiz question—is itself built by an AI pass over each lesson’s text and walkthrough transcripts, separately for Java and Kotlin, and stripped from the site at build time.
Several rules keep this defensible. A guard checks every model output before it is stored: it may name only sections from the menu, may contain no code, and may not mention the question’s name or anything from its text, and anything that fails is withheld rather than shown. Nothing is released before the quiz window has closed for everyone, nor within eight hours of the student’s own sitting ending, so nothing about a sitting can reach a student while a classmate may still be taking the quiz. Only official proctored sittings count, and staff never see feedback on their own accounts. Silence is a valid answer: in the first round, four in five sittings earned no pointer at all, and the page says so rather than inventing something to say. We read a sample of the first round’s output, with the attempts behind it, before releasing it, and every prompt and response is kept so any pointer can be traced back to what produced it.
Pitch Practice Feedback. Students preparing elevator pitches for their project ideas need a way to practice without requiring staff time for every attempt. They record themselves, get an automatic transcription, and receive encouraging feedback on clarity, delivery, and timing—letting them iterate before presenting to real people.
Submission Validation. Hundreds of project plan submissions need basic quality screening. An LLM validates that submitted plans are genuine implementation plans rather than placeholder text, providing instant feedback so students can fix issues immediately rather than waiting for manual review.
Where the Models Run. The AI endpoints used by the CS 124 courseware are provided by the University of Illinois and approved for use with student data. Transcription runs on our own servers. The coding agent students use for the project runs under their own accounts and is pointed at course content, not the other way around.