Logo
CASE STUDY / JANUARY 20, 2026 / MACHINE LEARNING / PYTHON / SCIKIT-LEARN / PYGAME / NEURAL NETWORKS

Chrome Dino AI Pilot

Artificial Intelligence capable of piloting the Chrome dinosaur completely autonomously, avoiding obstacles in real time via a neural network.

01 / BRIEF

Solution

Artificial Intelligence capable of piloting the Chrome dinosaur completely autonomously, avoiding obstacles in real time via a neural network.
02 / CONSTRAINTS

Problem

  • Critical class imbalance: continuous running (95%) vs obstacle jumping (5%).
  • Lazy model tendency to never jump to maximize its global accuracy.
  • Reducing real-time decision latency below the critical 1ms mark per prediction.
  • Reliable avoidance of various obstacles (single, multiple cacti and high altitude pterodactyls).
03 / SYSTEM

Architecture & Workflow

Real-time AI control architecture: Game state sensory inputs are extracted each frame, passed to an oversampled Multi-Layer Perceptron (MLP) neural classifier, and converted into sub-millisecond action commands inside the Pygame main loop.

Chrome Dino AI Pilot Architecture DiagramArchitecture & Workflow Diagram
01

Sensory Feature Extraction

Extracting distance, height, and game speed from current Pygame screen frame.

02

Oversampled Model Training

Training MLPClassifier neural network on rebalanced dataset to learn jump timing.

03

Sub-1ms Pygame Inference

Running serialized model inside game loop for real-time obstacle avoidance.

04

Evaluation & Performance Shift

Tracking confusion matrices, jump accuracy (85%), and dynamic speed feature weights.

1. Game State Sensing

Extracting 3 core input features from Pygame frame buffer: Distance to next obstacle, Obstacle height, and current Game speed.

2. Data Preprocessing & Oversampling

Rebalancing training dataset ratio (from 95:5 to balanced 50:50) using Random Oversampling on manual play recordings.

3. Neural Network Classifier (MLP)

Scikit-Learn MLPClassifier neural network trained to predict optimal action class (0: Run, 1: Jump).

4. Real-Time Inference Loop

Serialized .pkl model executing predictions in < 1ms per frame to trigger instant Pygame jump events.

04 / EVIDENCE

Results / impact

  • Jump accuracy increased from 40% to 85% thanks to oversampling.
  • 87% overall prediction accuracy across all obstacles.
  • Minimal inference latency below 1ms per decision.
  • Complete autonomous survival of the dinosaur on extended play sessions.

Avant Oversampling (Précision 66%)

90%
Course
10%
Saut
60%
FN (Crash)
40%
Saut OK

Le modèle "préfère" courir car c'est l'action ultra-majoritaire.

Après Oversampling (Précision 87%)

89%
Course
11%
Saut
15%
FN
85%
Saut OK

La diagonale est équilibrée : le modèle a appris l'importance vitale du saut.

Stratégie Qualité des Données : Optimisation par Oversampling

Description du Défi Technique

Lors des premières itérations, le modèle souffrait d'un déséquilibre majeur de classes : les phases de course représentaient plus de 95% du dataset. L'IA a donc développé une stratégie paresseuse en préférant ne jamais sauter pour maximiser statistiquement sa précision globale, au détriment de la survie (60% de collisions non évitées).

Analyse de la Matrice de Confusion

Version Initiale (66% de Précision) :

Un taux critique de Faux Négatifs (FN) est observé. Le modèle prédit "Course" lorsqu'un obstacle apparaît car il n'a pas appris la priorité absolue du saut.

Version Optimisée (87% de Précision) :

En appliquant une technique d'Oversampling, j'ai rééquilibré le poids de la classe minoritaire ("Saut").

Résultat : La diagonale de la matrice est désormais équilibrée. Le taux de succès des sauts est passé de 40% à 85%, prouvant que le modèle donne la priorité à l'évitement d'obstacle.

Feature Importance: Evolution of Decision Making

The AI's brain does not process information the same way depending on game intensity. This Radar Chart illustrates how the model adapts its priorities to survive at high speed.

Feature Importance

"The AI places 40% more importance on speed past 1000 points. At the start of the game, the model focuses almost exclusively on obstacle distance. However, as the game speeds up, the MLP shifts its attention to game speed: jump timing becomes more critical than the simple presence of the object."

DURATION3 weeks (R&D -> Simulation)
TEAM SIZESolo developer
ROLEReinforcement Learning Developer
TOOLS USEDPython · Pygame · NEAT Algorithm · Scikit-Learn
STACK

Tools used

Machine LearningPythonScikit-LearnPygameNeural Networks
RETROSPECTIVE

What I learned

  • Mastery of handling imbalanced datasets using Oversampling to rebalance the minority class ("Jump").
  • Optimization of real-time inference under the critical 1ms mark in Pygame.
  • Concrete use of the confusion matrix to analyze AI model performance.
NEXT ITERATION

What I would improve next

  • Training with Deep Reinforcement Learning (DQN) instead of the MLP Classifier.
  • Optimization of the model to adapt to faster dynamic speed changes.
REFERENCES

Links

GitHub RepositoryLINK
LinkedIn Demo & ArticleLINK
IMAGE GALLERY

Visual references

Inference Loop Architecture
01 / 03|Inference Loop Architecture
TECHNICAL DOCUMENTATION

Detailed technical documentation

01

Feature Engineering

Completed
Identification des variables critiques : distance de l'obstacle, hauteur, vitesse du jeu.
02

Data Collection

Completed
Capture de plus de 4 000 lignes de données en jouant manuellement pour l'entraînement.
03

Training

Completed
Mise en place d'un réseau de neurones (MLPClassifier) via Scikit-Learn.
04

Test & Analysis

Completed
Analyse des performances initiales (66% de précision sur les sauts).
05

Retraining & Optimization

Completed
Application des techniques de rééquilibrage (Oversampling) pour atteindre 87% de précision.
06

Inference

Completed
Déploiement du modèle .pkl dans la boucle Pygame pour la prise de décision en temps réel (< 1ms).