← Back to notes
Flutter & Cross-Platform 2026-10-10 00:23 13 min read Local copy

DecisionEngine: Teaching local AI to play Tetris

DecisionEngine: Teaching local AI to play Tetris
Jhin Lee for Google Developer Experts

Posted on Oct 10 • AI-assisted

DecisionEngine: Teaching local AI to play Tetris
#dart #flutter #ai #tutorial

Beyond chat with llamadart · Part 1

A piece is falling. There are a few places it could land. Somewhere in the background, a local AI model is choosing one.

If it takes too long, the game carries on without it.

Actual Laya Tetris browser run with the tuned choice player, 51 lines cleared and its candidate probabilities visible

The tuned player in an actual browser run: 51 lines cleared using WebGPU, Q8_0, and mixed candidates. This illustration is separate from the recorded benchmarks below.

Try the browser demo · Open the example

This opens Beyond chat with llamadart, a five-part series about cross-platform local AI for Dart and Flutter. We’ll explore decisions, embeddings, speech, images, and an application that connects them.

Tetris gives the experiment a visible consequence: the model picks a landing, and the board changes.

The game already has a useful heuristic. What does adding a decision model buy us? Tetris lets us compare different questions, training, and the cost of waiting for an answer.

Give the model a smaller job

The planner builds the options without AI. From the active piece’s current pose, it tries rotations, horizontal shifts, and a hard drop, rejecting blocked paths.

Each candidate is a Move: a reachable landing, the key sequence that gets there, and the resulting board measurements. Duplicate landings keep the shortest discovered sequence.

The game simulates each landing to count cleared lines, holes, stack height, and bumpiness. describeCompact() turns those consequences into the text sent to Laya.

The model chooses among descriptions of these planned moves. It does not generate the keys.

Schematic Tetris board with three candidate placements and an illustrative Laya choice of B

Three illustrative landings, with B shown as an example answer. Laya receives descriptions, not board images. Game code validates and executes the move.

Gravity keeps running. A target can become unreachable before the answer arrives, and the game records those failures.

Where Jev comes in

TypeSafe’s Jev gives this approach a useful vocabulary. TypeSafe calls it System One: provide a state and typed questions, then use the structured answers directly in code.

Laya is a separate open-weight model with a Jev-compatible interface. Llamadart follows Jev’s vocabulary and Laya’s request format; this experiment runs Laya locally.

The results here measure Laya and its Tetris-tuned head, not Jev.

Make your first decision with llamadart

A ModernBERT encoder runs through llama.cpp, and a decision head produces the answer. Each question takes an encoder pass without generating prose.

Game engine and model round trip

One question through the local pipeline. DecisionEngine prepares the tokens and decodes the model’s scores. The game continues running while it awaits the result.

ModernBERT reads the question, options, and state together. The head scores the positions marking each option; the decoder turns those scores into probabilities. Tuning changes the head while the encoder stays frozen.

Load a model and ask for a move

Add llamadart to a Dart project with dart pub add llamadart, save the following as bin/decision_walkthrough.dart, and run dart run bin/decision_walkthrough.dart on a supported native target. The documented decision path is validated on macOS; consult the support matrix for other targets. The first run downloads a roughly 421 MB encoder and 106 MB head. This standalone example prints one decision; it does not launch the Tetris game.

import 'package:llamadart/llamadart.dart';

Future<void> main(List<String> args) async {
  final useTunedHead = args.contains('--tuned');
  const repo = 'fr0stbit3/laya-gguf';
  const revision = 'ce2afdc0a8766af56a29a22dcf4a781e1f5c7d3c';

  final decisions = await DecisionEngine.load(
    DecisionModel(
      encoder: ModelSource.huggingFace(
        repoId: repo,
        revision: revision,
        filePath: 'laya-Q8_0.gguf',
      ),
      head: ModelSource.huggingFace(
        repoId: useTunedHead ? 'leehack/laya-tetris-head' : repo,
        revision: useTunedHead
            ? '465546a595ee2e8e3b212b8cb16829205d5dfab6'
            : revision,
        filePath: useTunedHead
            ? 'laya-head-tetris.safetensors'
            : 'laya-head.safetensors',
      ),
    ),
    params: const DecisionModelParams(
      device: ComputeDevice.auto,
      threads: 4,
    ),
  );

  try {
    // Illustrative descriptions, not a captured game turn.
    const candidates = <String, String>{
      'A': 'columns 1 to 2: clears zero lines, no new holes, '
          'stack 5 rows, bumpiness 8',
      'B': 'columns 5 to 6: clears one line, no new holes, '
          'stack 3 rows, bumpiness 4',
      'C': 'columns 9 to 10: clears zero lines, adds one hole, '
          'stack 6 rows, bumpiness 10',
    };

    final result = await decisions.systemOne(
      state: {'game': 'tetris', 'piece': 'O'},
      questions: {
        'move': DecisionQuestion.choice(
          'Which placement of the falling Tetris piece is best?',
          criteria: candidates,
        ),
      },
    );

    final ChoiceAnswer answer = result.choices['move']!;
    final String selectedDescription = candidates[answer.choice]!;

    print('Choice: ${answer.choice}');
    print('Probabilities: ${answer.probabilities}');
    print('Candidate: $selectedDescription');
  } finally {
    await decisions.dispose();
  }
}

DecisionModel identifies the encoder and head. load prepares their files and owns both resources. systemOne answers the named questions about one state; criteria maps each candidate label to its description.

Read the typed answer

result.choices['move'] returns a ChoiceAnswer. Its choice is the selected label, and probabilities contains every option’s probability. Here is an illustrative rendering of those fields, not captured output; your run may choose a different option.

{
  "choice": "B",
  "probabilities": {
    "A": 0.15,
    "B": 0.75,
    "C": 0.10
  }
}

The example uses candidates[answer.choice] to recover the selected description. In the game, that same lookup leads back to a Move with its landing and planned keys. The tuned player chooses from the probabilities itself so it can break ties randomly; ChoiceAnswer.choice uses the first highest-probability option. A probability of 0.75 is not a 75% guarantee of success.

From B to button presses

The mapping stays in game memory:

B → original candidate → Move {
  landing,
  planned keys,
  board measurements
}

Before pressing anything, the game checks that it is still controlling the same piece. Gravity has continued during inference, so it runs the planner again and finds a route to B’s landing from the current pose. It queues that route and presses one action at the configured interval.

If the piece has already locked, the answer is discarded. If the landing is now blocked, the game records the failure and hard-drops. The execution code handles those cases.

The API supports three kinds of question:

Kind Result Example
Choice A selected option and option probabilities Which landing should we use?
Score An expected value across ordered levels How urgent is this request?
Noul Probability of a yes/no statement Does this move clear a line?

“Noul” keeps the existing Jev and Laya vocabulary. A score can fall between levels, and a probability is not a boolean until the application applies a rule.

A chat model could also choose a move. There is no matched chat-model benchmark here, so this experiment cannot establish an advantage over LLMs.

Two settings to keep separate

Setting What it changes Options
Player How a move is selected Human, heuristic, random, Laya judge, checklist, base choice, tuned choice
Candidates Which planned moves are considered 3 best + 3 random, 6 best, 6 random, all generated legal moves

The heuristic player considers all generated legal moves. Base choice accepts at most six, so “all” falls back to the mixed shortlist for that player. Tuned choice handles larger sets through knockout rounds. The comparisons below use the mixed shortlist unless stated otherwise.

First result: one line

Give the base head six candidates and ask it to choose. In the short recorded benchmark, that player clears an average of one line. Random choice also clears one.

Asking “is this a good move?” about each candidate brings the mean to 18 lines.

A narrower checklist reaches 35 lines: ask whether a move clears a line and whether it creates holes, then subtract the second probability from the first.

The game already knows whether a move clears a line or creates a hole. The checklist tests how well Laya interprets those facts from text; an application could read the computed values directly. Six candidates can mean twelve model questions.

That improvement costs work. In this implementation, batching the questions into a call does not turn them into one encoder pass.

All three receive a shuffled shortlist: the heuristic’s three best moves plus three random remaining moves. Could training make the single-choice approach useful too?

Let the heuristic be the teacher

Only the decision head is tuned; the encoder stays frozen. Run dart run bin/decision_walkthrough.dart --tuned to try the pinned Tetris head with the same encoder. One sample decision is a smoke check; the gameplay comparison comes later.

The dataset has 16,000 training questions and 2,000 validation questions from other games. Labels come from the heuristic; tied best moves share the target probability.

Agreement with heuristic-best options on 2000 validation questions: random 27.5 percent, base 30.5 percent, published tuned head 75.7 percent

On the held-out questions, the base head scored 30.5% accuracy. A random pick scored 27.5%; the task can have multiple equally best options, so this is not simply one chance in six.

The published head reaches 75.7%. Eight documented 12-epoch runs ranged from 75.5% to 78.5%. The training recipe includes the full results.

This measures imitation of the heuristic, not optimal play. MPS training was not bit-for-bit repeatable.

A useful control: the base head still played roughly randomly when given the tuned question format. New wording alone did not explain the improvement.

Train a head of your own

Training happens in the Python notebook; DecisionEngine loads and runs the resulting head. From the prepared example/laya_tetris project, generate the seeded dataset:

dart run bin/make_dataset.dart dataset

Then run training/laya_head_tuning.ipynb using the documented Python environment. The notebook keeps the encoder frozen, trains the head, evaluates held-out questions, and writes laya-head-tetris.safetensors. Its embedded configuration lets llamadart load it without a separate config file.

In compare_heads.dart, replace the tunedHead: source with ModelSource.path('training/laya-head-tetris.safetensors'), adjusting the path to your exported file. For the first walkthrough, use the same source as DecisionModel.head. Keep the matching encoder; the Dart API runs the head after Python training.

Compare two heads on one encoder

For an app that switches between base and tuned players, load the encoder once and attach both heads. The essential calls are below. The complete runnable comparison includes the pinned model sources, request, main() invocation, and cleanup, including partial loading failures. Save it as bin/compare_heads.dart in the same Dart project and run dart run bin/compare_heads.dart.

// Excerpt from compare_heads.dart; setup and cleanup are in the file.
final engine = await LlamaEngine.load(
  LlamaModel(encoder),
  params: const DecisionModelParams().encoderModelParams,
);
final base = await DecisionEngine.attach(engine, head: baseHead);
final tuned = await DecisionEngine.attach(engine, head: tunedHead);

final baseResults = await base.systemOneBatch([request]);
final tunedResults = await tuned.systemOneBatch([request]);
print(baseResults.single.choices['move']!.probabilities);
print(tunedResults.single.choices['move']!.probabilities);

DecisionRequest packages a state and its questions for systemOneBatch. Results retain request order; .single reads this one-request comparison. For a larger comparison, supply several requests to each head. Each question still needs its own encoder pass. Attached heads borrow the encoder, so dispose the heads first, then the encoder. In a game, keep them loaded across turns.

Back to the board

The tuned choice player clears 48 lines and reaches the 150-piece cap in every game. Then the plain heuristic clears 56.

Historical mean lines cleared: random 1, base choice 1, base yes-no judge 18, base checklist 35, tuned choice 48 and heuristic baseline 56

Recorded turn-based results on M4 Max, Metal, Q8_0. Model and random rows use 3 best + 3 random candidates; the heuristic () considers all legal moves. These are descriptive means without per-seed intervals. Source: Tetris benchmark.*

The model learned something useful. The teacher still leads.

For six candidates, the tuned player needs just one question.

The heuristic considers all legal landings; the model’s shortlist already includes heuristic-selected moves. This is not a model learning Tetris unaided.

The board does not wait

In a live game, decision quality includes arriving on time. A stronger answer that arrives after the piece locks cannot help that move.

Jev’s documented hosted API involves a network request. Running locally removes that round trip and its variability. Whether it is faster overall still depends on the device, model, and workload; this experiment does not include a matched Jev benchmark.

That gives task-specific training a practical purpose. A small local model does not have to be the strongest general model to be useful. It needs to make sufficiently good decisions within the application’s deadline. Here, tuning improves the single-choice player while retaining the same encoder and decision-head architecture.

One six-option, 175-token question took about 20 milliseconds on Metal in the recorded M4 Max measurements.

Historical six-option question latency ranges on M4 Max: Metal 19.6 to 19.7 ms, CPU eight threads 83.5 to 85.2, six threads 108.4 to 110.8, four threads 157.9 to 159, two threads 306.3 to 312.5

Observed ranges across three runs per setting, using Q8_0 and a fresh engine for each setting. Model preparation is excluded; these are question timings, not full game-turn latency.

More candidates mean more work. With all legal moves, the tuned player runs knockout rounds in groups of up to six.

Across 40 documented real-time games, the tuned head averaged 52.3 lines with the mixed shortlist and 51.5 with all candidates. Every answer arrived before the piece locked.

These live games and the earlier capped tests have different conditions. Their scores belong to separate experiments.

Does clearing lines prove it works

It proves the player can clear lines. The integration needs other checks.

The validation record compares the implementation with Laya 0.3.5. A 24-question fixture showed no changed decisions across the measured configurations.

A broader 187-question set caught a Q8_0 CPU yes/no probability changing from 0.694 to 0.457, crossing the 0.5 threshold. F32 had no changed decisions on that set.

A smaller model representation needs its own evaluation on the questions your app asks.

Game tests cover rules, planning, candidate selection, and loading failures. Reference checks cover numerical agreement. Neither is replaced by a high game score.

Take the pattern into your app

Load once at application startup, reuse the model for decisions, and dispose it when the owning feature closes. Keep model preparation out of the game’s per-piece loop.

The decision-model guide covers typed keys, loading options, and runtime support when you move beyond this example.

Outside Tetris, the same pattern might choose a support category, route a request, or select among actions allowed by a simulation. Use clear descriptions and keep the available actions bounded by application state.

The current Laya format has a 512-token sequence limit and a smaller shared question-and-option budget. Large descriptions can be truncated. Confidence is derived from the returned distribution; it is not a calibrated guarantee of correctness.

Try changing the player

If the only goal were to build the strongest player in this comparison, the heuristic would be the obvious starting point.

The progression from one line to 48 makes local specialization worth exploring. Training improved this narrow task without requiring a larger inference architecture. The practical question is whether that quality is useful within the target device’s latency budget, with a simple heuristic beside it to keep the comparison honest.

Llamadart provides a way to explore that through Dart and Flutter, with a native llama.cpp path and a compatible WebGPU path. The measured M4 Max results should stay attached to that hardware and setup rather than becoming a claim about every phone or browser.

Try the Tetris demo, switch between the base and tuned players, and watch both the board and the decision panel. The failed moves are part of what makes the experiment worth looking at.

Next in the series: working with local embeddings.

Complete head comparison

Save this as bin/compare_heads.dart in your Dart project and run dart run bin/compare_heads.dart. The candidates are illustrative, as in the first walkthrough.

import 'package:llamadart/llamadart.dart';

Future<void> main() async {
  const repo = 'fr0stbit3/laya-gguf';
  const revision = 'ce2afdc0a8766af56a29a22dcf4a781e1f5c7d3c';
  await compareHeads(
    encoder: ModelSource.huggingFace(
      repoId: repo,
      revision: revision,
      filePath: 'laya-Q8_0.gguf',
    ),
    baseHead: ModelSource.huggingFace(
      repoId: repo,
      revision: revision,
      filePath: 'laya-head.safetensors',
    ),
    tunedHead: ModelSource.huggingFace(
      repoId: 'leehack/laya-tetris-head',
      revision: '465546a595ee2e8e3b212b8cb16829205d5dfab6',
      filePath: 'laya-head-tetris.safetensors',
    ),
    request: DecisionRequest(
      state: {'game': 'tetris', 'piece': 'O'},
      questions: {
        'move': DecisionQuestion.choice(
          'Which placement of the falling Tetris piece is best?',
          criteria: {
            'A': 'columns 1 to 2: clears zero lines, no new holes, stack 5 rows, bumpiness 8',
            'B': 'columns 5 to 6: clears one line, no new holes, stack 3 rows, bumpiness 4',
            'C': 'columns 9 to 10: clears zero lines, adds one hole, stack 6 rows, bumpiness 10',
          },
        ),
      },
    ),
  );
}

// Illustrative candidate descriptions, not a captured game turn.
Future<void> compareHeads({
  required ModelSource encoder,
  required ModelSource baseHead,
  required ModelSource tunedHead,
  required DecisionRequest request,
}) async {
  final engine = await LlamaEngine.load(
    LlamaModel(encoder),
    params: const DecisionModelParams().encoderModelParams,
  );

  DecisionEngine? base;
  DecisionEngine? tuned;
  try {
    base = await DecisionEngine.attach(engine, head: baseHead);
    tuned = await DecisionEngine.attach(engine, head: tunedHead);

    final baseResults = await base.systemOneBatch([request]);
    final tunedResults = await tuned.systemOneBatch([request]);

    print('Base: ${baseResults.single.choices['move']!.probabilities}');
    print('Tuned: ${tunedResults.single.choices['move']!.probabilities}');
  } finally {
    try {
      await tuned?.dispose();
    } finally {
      try {
        await base?.dispose();
      } finally {
        await engine.dispose();
      }
    }
  }
}

Top comments (0)

Subscribe

For further actions, you may consider blocking this person and/or reporting abuse