For facilitators · Running a cohort

THE RUBRIC

How to grade a game design course without ever grading "fun." This is the assessment engine from the MIT studio, rebuilt for anyone running FIRST PLAYABLE as a cohort, a club, or a self-paced run. If you are a solo learner, the last section turns all of it into a self-check.

The one rule that changes everything

You are not grading whether the game is good. You are grading whether the designer ran the loop: build, play, learn, change, repeat. A clumsy game that visibly improved across four playtests scores higher than a slick game that was thrown together the night before.

A great build on the final night cannot rescue a project that was never iterated along the way.

The core grading stance of MIT CMS.608, Philip Tan & Richard Eberhardt

/ what you collect

The four artifacts

Every build challenge produces the same four things to assess. Three come from the team, one from each individual. Plus a playable version brought to playtest days. Collect these and you can grade fairly without ever arguing about taste.

Team · 01

Rules & materials

Everything needed to play: the written rules and the component list. Graded for legibility. The real test is whether a stranger can play the game start to finish with no help from the designer in the room.

Team · 02

The changelog

Release notes for a game. After each design or playtest session, the team logs what they changed and why. This is the single best evidence of iteration and the heart of the grade.

Team · 03

The postmortem

A short presentation with visuals. The team walks through their process: what they tried, what broke, what they learned. Five minutes is plenty. Reflection over polish.

Individual · 04

Teamwork report

One page, written alone. How did you work inside the team? What was your role, where did you struggle, what would you do differently? This is how you grade individuals fairly on a group project.

/ how you score

The rubrics

Four levels for each thing you assess: Not yet Developing Solid Exemplary. Suggested weights are marked, but the point is the language, not the math. Use these descriptions out loud when you give feedback.

Rules & materials weight · 25%
CriterionWhat you are looking for, low to high
A stranger can play it
Not yet Needs the designer present to play at all.
Developing Playable but readers hit several questions the rules cannot answer.
Solid A new group gets through a full game with only minor confusion.
Exemplary A cold group plays correctly, first try, no questions.
Component list is exact Not yet "some pieces."   Developing mostly listed.   Solid precise counts and types.   Exemplary exact counts, clear naming, setup shown.
Steps are clear & ordered Not yet wall of text.   Developing steps exist but jump around.   Solid short active steps in order.   Exemplary ordered, active voice, exceptions flagged where they happen.
The changelog weight · 35% · the big one
CriterionWhat you are looking for, low to high
Evidence of real iteration
Not yet One or two entries, all from the last day.
Developing Regular entries but changes look cosmetic.
Solid Steady entries across the whole window with real design changes.
Exemplary A clear arc of change where you can watch the game find itself.
Changes tie to playtests Not yet changes seem random.   Developing some link to feedback.   Solid most changes name the problem they fix.   Exemplary every change cites what a player did that triggered it.
The "why" is there Not yet only lists what changed.   Developing occasional reasoning.   Solid reasons given for most changes.   Exemplary reasoning shows a growing theory of the game.
Postmortem & reflection weight · 25%
CriterionWhat you are looking for, low to high
Names what they learned
Not yet Describes the game, not the process.
Developing Lists events but not lessons.
Solid States clear lessons and how they got there.
Exemplary Lessons are specific, honest, and useful to other designers.
Owns the mistakes Not yet "it went great."   Developing vague regrets.   Solid names real missteps.   Exemplary dissects a failure and what it taught.
Individual teamwork report weight · 15%
CriterionWhat you are looking for, low to high
Honest self-account
Not yet Restates the group's work as their own.
Developing Generic "I helped with everything."
Solid Clear personal role and contributions.
Exemplary Role, friction, and growth all named with specifics.
/ the engine room

How to run a playtest day

Playtest days are where the grade actually gets earned. The whole method depends on running them well. Here is the protocol.

Swap, don't present

Teams trade games and play each other's. Designers do not pitch their game first. The rules sheet has to carry the whole load, because that is what gets graded.

Designers watch in silence

The designing team observes their game being played and writes down every moment a player hesitates, asks a question, or does something unexpected. They may not coach. Confusion is the data.

Debrief with behavior, not opinion

Feedback is about what happened, not whether it was fun. "Three players skipped the trading phase entirely" beats "I liked it." Train the room to report actions.

Log before you leave

Every team writes changelog entries that same session while the playtest is fresh. This single habit is what separates a real iteration trail from a fiction written on the last night.

Hold the deadlines

Due dates do not move. The lesson is that a game ships in the state it is in, and that the way to ship something good is to iterate early and often, not to crunch at the end. Moving a deadline quietly teaches the opposite of the entire course.

/ pacing

Two short builds, one long one

The shape of a cohort

  • Boss 1 & 2 are short. Run each over about three weeks: form teams, build, two or three playtest days, present.
  • Boss 3 is long. Give it roughly six weeks. It is the one with outside research and a real system to model, so it needs more iteration room.
  • Reshuffle teams between builds so people work with new collaborators and the teamwork report stays honest.
  • Two to four players per game keeps playtests fast and rules tight.

Running it solo or self-paced

  • You are your own facilitator. Use the rubric language above on yourself, gently and honestly.
  • Find three playtesters. Friends, family, a Discord. The "stranger can play it" test still works with one person who has never seen your game.
  • Keep the changelog anyway. Even solo, log what you changed and why after every test. It is the most valuable artifact you will make.
  • Set fake deadlines and keep them. The discipline is the lesson, with or without a grade.