How to grade a game design course without ever grading "fun." This is the assessment engine from the MIT studio, rebuilt for anyone running FIRST PLAYABLE as a cohort, a club, or a self-paced run. If you are a solo learner, the last section turns all of it into a self-check.
You are not grading whether the game is good. You are grading whether the designer ran the loop: build, play, learn, change, repeat. A clumsy game that visibly improved across four playtests scores higher than a slick game that was thrown together the night before.
A great build on the final night cannot rescue a project that was never iterated along the way.
The core grading stance of MIT CMS.608, Philip Tan & Richard Eberhardt
Every build challenge produces the same four things to assess. Three come from the team, one from each individual. Plus a playable version brought to playtest days. Collect these and you can grade fairly without ever arguing about taste.
Everything needed to play: the written rules and the component list. Graded for legibility. The real test is whether a stranger can play the game start to finish with no help from the designer in the room.
Release notes for a game. After each design or playtest session, the team logs what they changed and why. This is the single best evidence of iteration and the heart of the grade.
A short presentation with visuals. The team walks through their process: what they tried, what broke, what they learned. Five minutes is plenty. Reflection over polish.
One page, written alone. How did you work inside the team? What was your role, where did you struggle, what would you do differently? This is how you grade individuals fairly on a group project.
Four levels for each thing you assess: Not yet Developing Solid Exemplary. Suggested weights are marked, but the point is the language, not the math. Use these descriptions out loud when you give feedback.
| Criterion | What you are looking for, low to high |
|---|---|
| A stranger can play it |
Not yet Needs the designer present to play at all.
Developing Playable but readers hit several questions the rules cannot answer.
Solid A new group gets through a full game with only minor confusion.
Exemplary A cold group plays correctly, first try, no questions.
|
| Component list is exact | Not yet "some pieces." Developing mostly listed. Solid precise counts and types. Exemplary exact counts, clear naming, setup shown. |
| Steps are clear & ordered | Not yet wall of text. Developing steps exist but jump around. Solid short active steps in order. Exemplary ordered, active voice, exceptions flagged where they happen. |
| Criterion | What you are looking for, low to high |
|---|---|
| Evidence of real iteration |
Not yet One or two entries, all from the last day.
Developing Regular entries but changes look cosmetic.
Solid Steady entries across the whole window with real design changes.
Exemplary A clear arc of change where you can watch the game find itself.
|
| Changes tie to playtests | Not yet changes seem random. Developing some link to feedback. Solid most changes name the problem they fix. Exemplary every change cites what a player did that triggered it. |
| The "why" is there | Not yet only lists what changed. Developing occasional reasoning. Solid reasons given for most changes. Exemplary reasoning shows a growing theory of the game. |
| Criterion | What you are looking for, low to high |
|---|---|
| Names what they learned |
Not yet Describes the game, not the process.
Developing Lists events but not lessons.
Solid States clear lessons and how they got there.
Exemplary Lessons are specific, honest, and useful to other designers.
|
| Owns the mistakes | Not yet "it went great." Developing vague regrets. Solid names real missteps. Exemplary dissects a failure and what it taught. |
| Criterion | What you are looking for, low to high |
|---|---|
| Honest self-account |
Not yet Restates the group's work as their own.
Developing Generic "I helped with everything."
Solid Clear personal role and contributions.
Exemplary Role, friction, and growth all named with specifics.
|
Playtest days are where the grade actually gets earned. The whole method depends on running them well. Here is the protocol.
Teams trade games and play each other's. Designers do not pitch their game first. The rules sheet has to carry the whole load, because that is what gets graded.
The designing team observes their game being played and writes down every moment a player hesitates, asks a question, or does something unexpected. They may not coach. Confusion is the data.
Feedback is about what happened, not whether it was fun. "Three players skipped the trading phase entirely" beats "I liked it." Train the room to report actions.
Every team writes changelog entries that same session while the playtest is fresh. This single habit is what separates a real iteration trail from a fiction written on the last night.
Due dates do not move. The lesson is that a game ships in the state it is in, and that the way to ship something good is to iterate early and often, not to crunch at the end. Moving a deadline quietly teaches the opposite of the entire course.