Samik Hafeez
All work

Year 2 team project, then a 2026 rebuild I lead

OrOgins: getting a game and its AI to work together

OrOgins is an elemental strategy board game whose digital version was created in 2015. The brief was to modernise its interface and let people play against an AI. In my Year 2 project the game and the AI were built separately and never connected. Since June 2026 I have led a rebuild as project manager and technical lead, and the game and its AI now work together. The next step, under way, is removing the Python server so the game runs offline.

Context
Year 2 university project, then a 2026 rebuild with client reviews
My role
Year 2: Project Manager, Unity development and testing · Rebuild: project manager and technical lead
Team
Year 2: 3 students · Rebuild: a small team with a supervisor, weekly progress meetings and client reviews
Dates
Nov 2024 – Apr 2025 · June 2026 – present
Status
Playable against three AI levels; still needs a local Python server; offline C# version in progress

Stack

  • Unity 6
  • C#
  • Python
  • NumPy
  • Flask
  • pytest
  • Stable-Baselines3 (Year 2)

Scope of this pageThe rebuild is ongoing work with a client. Its figures come from the project's own test runs and recorded benchmarks (September 2026); the 80 tests were re-run, and passed, when this page was written.

The rebuilt OrOgins game running in Unity, Player vs AI on Hard. A 10 × 8 board: the Evolutionist side's two ape royals and element discs along the top row, the Creationist side's two human royals and element discs along the bottom, and trails painted yellow, blue, red and green where elements have moved. A narrow side panel shows the mode and difficulty buttons and “AI replied. Turn: Player1”.
ScreenshotThe rebuilt game against the Hard AI, September 2026: my Year 2 art, driven by the new engine. The side panel reads “AI replied. Turn: Player1 · You are Player1 (AI: learned)”.

The short version

Each point is expanded, with diagrams, in the sections below.

01Problem and constraints
A 2015 digital board game needed a modern interface and an AI opponent. The AI had to follow the game's real rules and become part of the game, not a separate program.
02My responsibility
Year 2: Project Manager and Scrum Master, Unity developer and tester. Rebuild: project manager and technical lead.
03The design
Unity only draws the board and sends clicks; Python owns the rules and the AI. The accepted next design moves both into C# inside Unity and keeps Python as the reference and training environment.
04What I implemented
Year 2: the Unity board, pieces, movement scripts and tests. Rebuild: the 10 × 8 engine, the 12-weight Hard AI and its move rules, my Unity game as the interface, and the research into removing Flask.
05The team’s part
Year 2: a teammate built the Gym environment and trained PPO. The rebuild started from an earlier 8 × 8 student version, with a supervisor and client reviews along the way.
06What was delivered
Year 2: a Unity game and a separate learning prototype. Rebuild: a playable game against Easy, Normal and Hard opponents, 80 passing tests, and an accepted plan for offline play.
07What testing established
The Hard AI wins 85% of 40 games against Easy and 57.5% against Normal. A fix cut its back-and-forth moves from 53.7% to 4.3% with no measurable loss of strength.
08Learned, and unfinished
Connect the halves first, build on the real rules before training anything, and prefer a model the team can explain. The game still needs Flask; the C# version is not finished.
On this page

The game, and the brief

OrOgins is a two-player elemental strategy game whose digital version was created in 2015. The brief, in my Year 2 project and again in the rebuild, was to modernise its interface and let people play it against an AI opponent.

Each side has two royals and eight element pieces: earth, water, fire and air, each beating the next in a cycle. Elements move like chess queens and paint the squares they cross. Royals can only travel forward or sideways along painted, empty squares, so the elements have to build their roads. A player wins by getting both royals to the rank in front of the opponent's home row. An element that crosses a square painted in the colour it beats wipes it and knocks off a royal standing there, and a player who loses a royal can no longer win.

Explanatory diagramDerived from env.py reset() and Game.cs (same piece order)

The board both halves of the project had to agree on

OrOgins starting position on the 10 × 8 board012345678901234567EaWaFiAiWoMaAiFiWaEaEaWaFiAiWoMaAiFiWaEaROW 0Creationist piecesROW 1Evolutionist goalROW 6Creationist goal (“Day 6”)ROW 7Evolutionist piecesEa earth · Wa water · Fi fire · Ai air · Wo woman · Ma man. Strength cycle: earth › water › fire › air › earth.
OrOgins starting position on the 10 × 8 board012345678901234567EaWaFiAiWoMaAiFiWaEaEaWaFiAiWoMaAiFiWaEaROW 0Creationist piecesROW 1Evolutionist goalROW 6Creationist goal (“Day 6”)ROW 7Evolutionist piecesEa earth · Wa water · Fi fire · Ai air · Wo woman · Ma man.Strength cycle: earth › water › fire › air › earth.
The 10-column, 8-row starting position used by the Python environment. Each side must bring its man and woman to the opposite goal row; elements convert neutral squares and capture weaker elements.

The board both halves of the project had to agree on

100%
OrOgins starting position on the 10 × 8 board012345678901234567EaWaFiAiWoMaAiFiWaEaEaWaFiAiWoMaAiFiWaEaROW 0Creationist piecesROW 1Evolutionist goalROW 6Creationist goal (“Day 6”)ROW 7Evolutionist piecesEa earth · Wa water · Fi fire · Ai air · Wo woman · Ma man. Strength cycle: earth › water › fire › air › earth.
Text version of this diagram

Ten columns by eight rows. Row 0 holds the Creationist pieces and row 7 the Evolutionist pieces, each in the order earth, water, fire, air, woman, man, air, fire, water, earth. Creationists aim for row 6 and Evolutionists for row 1. Element strength runs earth over water, water over fire, fire over air and air over earth.

· Phase 1 · 2024–25

Year 2: two halves that never met

I managed our three-person team and built the game in Unity; a teammate encoded the rules in a Python Gym environment and trained a PPO agent on it. The plan was to export the trained policy and run it inside Unity. The export was never written, nothing in Unity called the model wrapper, and when the module ended the game and the AI still could not talk to each other.

Year 2: who did what, from the design work, reflection and contribution tables
AreaMy partTeammates
ManagementProject Manager; Scrum Master for Sprints 1 and 3; project plan in Microsoft Project; risk assessment and quality plan.Trello board, GitHub; Scrum Master for Sprint 2.
Requirements and designFunctional and non-functional requirements; chose Unity over Unreal; board and piece art.Five-phase plan and player-experience questions.
Unity gameUnity set-up, board, piece prefabs, and the C# scripts for pieces, moves, capture and neutral squares.Unity game logic assigned to a teammate; the element rules stayed unfinished.
Learning environmentHelped with the environment design and debugging.Gym environment, state and action spaces, and the Pygame prototype rules.
ModelModel testing, debugging and monitoring training, with the teammate.Chose and trained PPO with Stable-Baselines3.
TestingWrote, ran and documented the test and validation scripts; analysed the failures.Model validation.
Explanatory diagramDerived from the academic repository (C# scripts, env.py, train.py, AIModelManager.cs) and the reports

Two systems that were meant to meet

Two systems that were meant to meetWhat existed on each side, and the planned bridge between them. Dashed blue parts were planned and never built; dashed red parts were started but unfinished.UNITY GAME · C# (MY BUILD)PLANNED INTEGRATIONPYTHON RL · TEAMMATE'S ENVIRONMENT AND TRAININGGame.cs controllercreates 10 pieces a side; neutralsquares fill the restPieces + move platesOrOrginsMan.cs, MovePlate.cs:click to move, captureNeutral-square ruleexperiment: fire pieces onlyElement, capture, victory andturn rulesnot implemented in UnityAIModelManagerBarracuda wrapper; never calledExport policy to ONNXMap Unity board → 100-valueobservationOriginsEnv (Gym)board rules, rewards, game-overchecksPPO trainingStable-Baselines3 MlpPolicy,100k steps configuredppo_origins.zipsaved SB3 modelPygame play loopopponent plays random legalmovesTest + validation scripts9 tests, 9 validators, 4 evaluatorsobservations, rewardschosen moveexercise the rules
Two systems that were meant to meetWhat existed on each side, and the planned bridge between them. Dashed blue parts were planned and never built; dashed red parts were started but unfinished.UNITY GAME · C# (MY BUILD)PLANNEDPYTHON RL · TEAMMATE'S ENVIRONMENT AND TRAININGGame.cs controllercreates 10 pieces a side;neutral squares fill therestPieces + move platesOrOrginsMan.cs,MovePlate.cs: click tomove, captureNeutral-square ruleexperiment: fire piecesonlyElement, capture,victory and turn rulesnot implemented in UnityAIModelManagerBarracuda wrapper; nevercalledExport policy to ONNXMap Unity board →100-value observationOriginsEnv (Gym)board rules, rewards,game-over checksPPO trainingStable-Baselines3MlpPolicy, 100k stepsconfiguredppo_origins.zipsaved SB3 modelPygame play loopopponent plays randomlegal movesTest + validationscripts9 tests, 9 validators, 4evaluators
  • Implemented
  • My contribution
  • Stored data
  • Incomplete, or a gap found in testing
  • Planned or proposed — not built
What existed on each side, and the planned bridge between them. Dashed blue parts were planned and never built; dashed red parts were started but unfinished.

The trained policy was saved in Stable-Baselines3's own format. Nothing exported it to ONNX, nothing in Unity called the Barracuda wrapper, and nothing mapped the Unity board to the environment's 100-value observation. That missing bridge is the project's main unfinished piece.

Two systems that were meant to meet

100%
Two systems that were meant to meetUNITY GAME · C# (MY BUILD)PLANNED INTEGRATIONPYTHON RL · TEAMMATE'S ENVIRONMENT AND TRAININGGame.cs controllercreates 10 pieces a side; neutralsquares fill the restPieces + move platesOrOrginsMan.cs, MovePlate.cs:click to move, captureNeutral-square ruleexperiment: fire pieces onlyElement, capture, victory andturn rulesnot implemented in UnityAIModelManagerBarracuda wrapper; never calledExport policy to ONNXMap Unity board → 100-valueobservationOriginsEnv (Gym)board rules, rewards, game-overchecksPPO trainingStable-Baselines3 MlpPolicy,100k steps configuredppo_origins.zipsaved SB3 modelPygame play loopopponent plays random legalmovesTest + validation scripts9 tests, 9 validators, 4 evaluatorsobservations, rewardschosen moveexercise the rules
Text version of this diagram

Parts

  • Unity game · C# (my build)
    • Game.cs controller — creates 10 pieces a side; neutral squares fill the rest; my contribution
    • Pieces + move plates — OrOrginsMan.cs, MovePlate.cs: click to move, capture; my contribution
    • Neutral-square rule — experiment: fire pieces only; incomplete or failed in testing; my contribution
    • Element, capture, victory and turn rules — not implemented in Unity; incomplete or failed in testing
    • AIModelManager — Barracuda wrapper; never called; incomplete or failed in testing
  • Planned integration (planned, not built)
    • Export policy to ONNX — planned, not built
    • Map Unity board → 100-value observation — planned, not built
  • Python RL · teammate's environment and training
    • OriginsEnv (Gym) — board rules, rewards, game-over checks; teammate
    • PPO training — Stable-Baselines3 MlpPolicy, 100k steps configured; teammate
    • ppo_origins.zip — saved SB3 model
    • Pygame play loop — opponent plays random legal moves; teammate
  • Test + validation scripts — 9 tests, 9 validators, 4 evaluators; my contribution; I wrote, ran and documented the model tests

Connections

  • Game.cs controller → Pieces + move plates
  • Pieces + move plates → Neutral-square rule
  • Game.cs controller → Element, capture, victory and turn rules (incomplete or failed in testing)
  • OriginsEnv (Gym) → PPO training: observations, rewards
  • PPO training → ppo_origins.zip
  • OriginsEnv (Gym) → Pygame play loop
  • ppo_origins.zip → Export policy to ONNX (planned, not built)
  • Export policy to ONNX → AIModelManager (planned, not built)
  • Map Unity board → 100-value observation → AIModelManager (planned, not built)
  • AIModelManager → Game.cs controller: chosen move (planned, not built)
  • Test + validation scripts → OriginsEnv (Gym): exercise the rules
Year 2 in detail: what the agent learned from, what the tests showed, and how we ran it

Reading the environment code shows how narrow the agent's choice really was. It picks a square; the environment then moves that piece to the first legal destination it finds. Rewards say “legal move” or “illegal move” and ±100 at the end, so good play was never directly rewarded.

Explanatory diagramDerived from OriginsEnv.step() and reset() in env.py and the PPO settings in train.py

One step of the observation–action–reward loop, as coded

One step of the observation–action–reward loop, as codedOne step of training as the environment implements it. Dashed red boxes mark mismatches found by reading the code: reward sign and declared observation bounds.PPO policy (MlpPolicy)one policy plays both sidesAction: a square, 0–79the policy never picks a destinationOwn piece on that square?Reward −1same player tries againAny legal move?Reward −1turn passesTake the first legal destinationfixed direction orderMove, convert path squares,capturereward +1Game over?both arrived, a man or womanlost, or no movesReward ±100 and resetsign follows the mover, not the winnerSwitch turnObservation: 100 integers80 squares, turn, 4 arrival flags,padding; bounds declared ±3 butpieces use ±6noyesnoyesyesnonext step
One step of the observation–action–reward loop, as codedOne step of training as the environment implements it. Dashed red boxes mark mismatches found by reading the code: reward sign and declared observation bounds.PPO policy (MlpPolicy)one policy plays bothsidesAction: a square, 0–79the policy never picks adestinationOwn piece on thatsquare?Reward −1same player tries againAny legal move?Reward −1turn passesTake the first legaldestinationfixed direction orderMove, convert pathsquares, capturereward +1Game over?both arrived, a manor woman lost, or nomovesReward ±100 andresetsign follows the mover,not the winnerSwitch turnObservation: 100integers80 squares, turn, 4 arrivalflags, padding; boundsdeclared ±3 but piecesuse ±6noyesnoyesyesno
  • Implemented
  • Decision
  • Incomplete, or a gap found in testing
One step of training as the environment implements it. Dashed red boxes mark mismatches found by reading the code: reward sign and declared observation bounds.

One step of the observation–action–reward loop, as coded

100%
One step of the observation–action–reward loop, as codedPPO policy (MlpPolicy)one policy plays both sidesAction: a square, 0–79the policy never picks a destinationOwn piece on that square?Reward −1same player tries againAny legal move?Reward −1turn passesTake the first legal destinationfixed direction orderMove, convert path squares,capturereward +1Game over?both arrived, a man or womanlost, or no movesReward ±100 and resetsign follows the mover, not the winnerSwitch turnObservation: 100 integers80 squares, turn, 4 arrival flags,padding; bounds declared ±3 butpieces use ±6noyesnoyesyesnonext step
Text version of this diagram

Parts

  • PPO policy (MlpPolicy) — one policy plays both sides; teammate
  • Action: a square, 0–79 — the policy never picks a destination
  • Own piece on that square?
  • Reward −1 — same player tries again
  • Any legal move?
  • Reward −1 — turn passes
  • Take the first legal destination — fixed direction order
  • Move, convert path squares, capture — reward +1
  • Game over? — both arrived, a man or woman lost, or no moves
  • Reward ±100 and reset — sign follows the mover, not the winner; incomplete or failed in testing
  • Switch turn
  • Observation: 100 integers — 80 squares, turn, 4 arrival flags, padding; bounds declared ±3 but pieces use ±6; incomplete or failed in testing

Connections

  • PPO policy (MlpPolicy) → Action: a square, 0–79
  • Action: a square, 0–79 → Own piece on that square?
  • Own piece on that square? → Reward −1: no
  • Own piece on that square? → Any legal move?: yes
  • Any legal move? → Reward −1: no
  • Any legal move? → Take the first legal destination: yes
  • Take the first legal destination → Move, convert path squares, capture
  • Move, convert path squares, capture → Game over?
  • Game over? → Reward ±100 and reset: yes
  • Game over? → Switch turn: no
  • Switch turn → Observation: 100 integers
  • Reward −1 → Observation: 100 integers
  • Reward −1 → Observation: 100 integers
  • Observation: 100 integers → PPO policy (MlpPolicy): next step

Without capture and victory rules in place, a win rate could not measure anything, so I tested what could be checked deterministically. The failures were the useful part: they pointed at missing rules rather than at the learning code.

Test and validation results recorded in the final reflection

  • Board dimensions Passed10 columns × 8 rows
  • Board initialisation MixedPieces and starting positionsTests passed; the stricter validator failed 1 of 4 checks.
  • Observation encoding PassedLength and piece codes of the 100-value observation
  • Element interactions MixedThe strength cycle, including fire against airTests passed; the air-against-earth validation crashed with an index error.
  • Movement FailedPieces move in their allowed directions“Evolutionist man should move up”.
  • Turn switching FailedThe turn passes after a move
  • Destination timing MixedThe goal rows are recognisedStalemate after a maximum number of turns failed.
  • AI decision checks MixedCapture and protective moves are availableChecks the rules, never the trained policy; strategic progress failed.
  • Whole-game validation Mixed26 rule checks across the environment15 of 26 passed.

Several copies of the environment existed, so these results describe the build tested at the time, not a single final version.

Terminal output: board dimensions test passed (8x10 grid confirmed), initial piece placement test passed, starting positions test passed, all board initialisation tests passed.
Test outputBoard initialisation tests passing, from the Year 2 final reflection.Open full size

We worked in Agile sprints with a Trello backlog, stand-ups and reviews; SDLC structured the game track and CRISP-DM the learning track. Sprint 1 ran long because Unity and the rules took far more time than planned, so by the time the environment and PPO arrived there was little room left to connect them.

Explanatory diagramDerived from the sprint plan and contribution tables in the final reflection

Three sprints, two tracks, one unfinished connection

Three sprints, two tracks, one unfinished connectionDecJan 2025FebMarAprSPRINTSSprint 1SM: meSprint 2SM: teammateS3GAME TRACK (SDLC)Requirements, plans, Unity board and piecesLEARNING TRACK(CRISP-DM)Gym environment+ PPO (teammate)INTEGRATIONTrained policy insideUnity: not reachedTESTINGTests + validation
Sprint goals and my part in each
SprintTeam goal (sprint plan)My part
Sprint 125 Nov – 16 MarInitialise the project, set up Unity, build the board and pieces.Scrum Master. Requirements, project, risk and quality plans; Unity setup, board, pieces and C# scripts.
Sprint 217 Mar – 3 AprCustom Gym environment, PPO training and integration.Model testing and debugging with the teammate who built the environment.
Sprint 34 Apr – 6 AprCompletion, evaluation, validation and the presentation.Scrum Master. Ran and documented the test and validation scripts; sprint review.
  1. Sprint 125 Nov – 16 MarScrum Master: me

    Initialise the project, set up Unity, build the board and pieces.

    My part Requirements, project, risk and quality plans; Unity setup, board, pieces and C# scripts.

    • Game track (SDLC) Requirements, plans, Unity board and pieces
  2. Sprint 217 Mar – 3 Apr

    Custom Gym environment, PPO training and integration.

    My part Model testing and debugging with the teammate who built the environment.

    • Learning track (CRISP-DM) Gym environment + PPO (teammate)
    • Integration Trained policy inside Unity: not reached
  3. Sprint 34 Apr – 6 AprScrum Master: me

    Completion, evaluation, validation and the presentation.

    My part Ran and documented the test and validation scripts; sprint review.

    • Integration Trained policy inside Unity: not reached
    • Testing Tests + validation
  • Sprint or work
  • Mine
  • Planned, not reached
Sprint dates from the final reflection. Tinted sprints are the ones I ran as Scrum Master; the dashed blue bar is the integration the team planned but did not reach.

Three sprints, two tracks, one unfinished connection

100%
Three sprints, two tracks, one unfinished connectionDecJan 2025FebMarAprSPRINTSSprint 1SM: meSprint 2SM: teammateS3GAME TRACK (SDLC)Requirements, plans, Unity board and piecesLEARNING TRACK(CRISP-DM)Gym environment+ PPO (teammate)INTEGRATIONTrained policy insideUnity: not reachedTESTINGTests + validation
Text version of this diagram
  • Sprint 1 (25 Nov – 16 Mar): Initialise the project, set up Unity, build the board and pieces. I was Scrum Master. My part: Requirements, project, risk and quality plans; Unity setup, board, pieces and C# scripts.
  • Sprint 2 (17 Mar – 3 Apr): Custom Gym environment, PPO training and integration. My part: Model testing and debugging with the teammate who built the environment.
  • Sprint 3 (4 Apr – 6 Apr): Completion, evaluation, validation and the presentation. I was Scrum Master. My part: Ran and documented the test and validation scripts; sprint review.
  • Game track (SDLC): Requirements, plans, Unity board and pieces (25 Nov – 16 Mar)
  • Learning track (CRISP-DM): Gym environment + PPO (teammate) (17 Mar – 3 Apr)
  • Integration: Trained policy inside Unity: not reached (17 Mar – 6 Apr)
  • Testing: Tests + validation (4 Apr – 6 Apr)
Explanatory diagramDerived from the CRISP-DM figure and text in the final reflection

CRISP-DM for the learning track

CRISP-DM for the learning trackCRISP-DMreinforcement learning has no fixed dataset: theenvironment is the data01Business understandingAn AI opponent for atwo-player strategy game02Data understandingRules, board state, actions andobservations03Data preparationRules coded into OriginsEnv;only legal moves allowed04ModellingPPO with Stable-Baselines3(teammate)05EvaluationMy test and validation scripts;rule checks, not a win rate06DeploymentPolicy inside the Unity game:not reached
  1. Business understandingAn AI opponent for a two-player strategy game
  2. Data understandingRules, board state, actions and observations
  3. Data preparationRules coded into OriginsEnv; only legal moves allowed
  4. ModellingPPO with Stable-Baselines3 (teammate)
  5. EvaluationMy test and validation scripts; rule checks, not a win rate
  6. DeploymentPolicy inside the Unity game: not reached
  • Stage
  • My contribution
  • Incomplete, or a gap found in testing
  • Not reached
The learning track mapped to CRISP-DM. Evaluation was mine and only partial; deployment into Unity was not reached.

CRISP-DM for the learning track

100%
CRISP-DM for the learning trackCRISP-DMreinforcement learning has no fixed dataset: theenvironment is the data01Business understandingAn AI opponent for atwo-player strategy game02Data understandingRules, board state, actions andobservations03Data preparationRules coded into OriginsEnv;only legal moves allowed04ModellingPPO with Stable-Baselines3(teammate)05EvaluationMy test and validation scripts;rule checks, not a win rate06DeploymentPolicy inside the Unity game:not reached
Text version of this diagram
  1. Business understanding: An AI opponent for a two-player strategy game
  2. Data understanding: Rules, board state, actions and observations
  3. Data preparation: Rules coded into OriginsEnv; only legal moves allowed
  4. Modelling: PPO with Stable-Baselines3 (teammate)
  5. Evaluation: My test and validation scripts; rule checks, not a win rate (my contribution)
  6. Deployment: Policy inside the Unity game: not reached (not done)
Situation
An unfamiliar engine, complex rules and two technologies developed side by side.
Task
Plan the work, build the Unity foundation and keep the team heading for a working game with an AI opponent.
Action
I set up the plans, built the board and scripts, and turned testing into a record of which rules worked and which did not.
Result
A Unity prototype and a separate trained-agent pipeline, but no working connection between them.
Learning
Plan smaller end-to-end milestones, agree one environment specification, and test the integration boundary first. It is exactly where the rebuild started.

· Phase 2 · June 2026 – present

The rebuild: one engine, and my game as its interface

In June 2026 I came back to OrOgins as project manager and technical lead of a rebuild, with a supervisor, weekly progress meetings and reviews with the client. The goal was the one Year 2 missed: a game you can actually play against the AI.

The starting point was an earlier 8 × 8 student version of the Python engine. The original game is played on 10 columns by 8 rows, so I reworked the engine to implement the original rules on the correct board: royals riding painted trails, captures by wiping tiles, and the win and can't-win conditions. Every rule became an automated test.

My Year 2 Unity game became the interface. It now holds no rules and no AI: it draws whatever state the Python engine sends back, and passes on the player's clicks. Players choose PvP or PvAI and Easy, Normal or Hard, and for the first time the game, in its proper interface, plays against a trained AI.

Explanatory diagramDerived from server_v1.py, src/ and the Unity frontend scripts (OrOginsApiClient.cs, OrOginsGameController.cs) in the rebuild

How the rebuilt game runs today

How the rebuilt game runs todayUnity draws the board and sends clicks; the Python engine owns every rule and the AI. Selecting a piece and committing a move are two HTTP calls. The dashed red box is the one thing still in the way.UNITY 6 · MY YEAR 2 GAME, NOW THE INTERFACEPYTHON PROCESS · MUST BE STARTED BEFORE THE GAMEMenu and status panelPvP or PvAI; Easy, Normal orHard; New GameBoard, tiles and piecesdraws the game state it is sent, inmy Year 2 artAPI clientno rules and no AI inside UnityFlask APIfive endpoints: new_game, state,legal_moves, move, healthRules engine10 × 8, the original rules; 80automated testsThree AI levelsEasy random · Normal greedy ·Hard learnedHard AI weights12 numbers in a 677-byte fileThe remaining dependencya second process has to berunning: no offline play, nobrowser demo, no mobilenew gameclicksHTTPAI replytries moves
How the rebuilt game runs todayUnity draws the board and sends clicks; the Python engine owns every rule and the AI. Selecting a piece and committing a move are two HTTP calls. The dashed red box is the one thing still in the way.UNITY 6 · INTERFACEPYTHON PROCESSMenu and statuspanelPvP or PvAI; Easy,Normal or Hard; NewGameBoard, tiles andpiecesdraws the game state it issent, in my Year 2 artAPI clientno rules and no AI insideUnityFlask APIfive endpoints:new_game, state,legal_moves, move,healthRules engine10 × 8, the original rules;80 automated testsThree AI levelsEasy random · Normalgreedy · Hard learnedHard AI weights12 numbers in a677-byte fileThe remainingdependencya second process has tobe running: no offlineplay, no browser demo,no mobileHTTP
  • Implemented
  • Stored data
  • Incomplete, or a gap found in testing
Unity draws the board and sends clicks; the Python engine owns every rule and the AI. Selecting a piece and committing a move are two HTTP calls. The dashed red box is the one thing still in the way.

This split is what finally made the game and the AI work together, and it is also why the next step is possible: the Unity side already holds no rules, so the server can later be swapped for local C# code without touching what players see.

How the rebuilt game runs today

100%
How the rebuilt game runs todayUNITY 6 · MY YEAR 2 GAME, NOW THE INTERFACEPYTHON PROCESS · MUST BE STARTED BEFORE THE GAMEMenu and status panelPvP or PvAI; Easy, Normal orHard; New GameBoard, tiles and piecesdraws the game state it is sent, inmy Year 2 artAPI clientno rules and no AI inside UnityFlask APIfive endpoints: new_game, state,legal_moves, move, healthRules engine10 × 8, the original rules; 80automated testsThree AI levelsEasy random · Normal greedy ·Hard learnedHard AI weights12 numbers in a 677-byte fileThe remaining dependencya second process has to berunning: no offline play, nobrowser demo, no mobilenew gameclicksHTTPAI replytries moves
Text version of this diagram

Parts

  • Unity 6 · my Year 2 game, now the interface
    • Menu and status panel — PvP or PvAI; Easy, Normal or Hard; New Game
    • Board, tiles and pieces — draws the game state it is sent, in my Year 2 art
    • API client — no rules and no AI inside Unity
  • Python process · must be started before the game
    • Flask API — five endpoints: new_game, state, legal_moves, move, health
    • Rules engine — 10 × 8, the original rules; 80 automated tests
    • Three AI levels — Easy random · Normal greedy · Hard learned
    • Hard AI weights — 12 numbers in a 677-byte file
  • The remaining dependency — a second process has to be running: no offline play, no browser demo, no mobile; incomplete or failed in testing

Connections

  • Menu and status panel → API client: new game
  • Board, tiles and pieces → API client: clicks
  • API client → Flask API: HTTP
  • Flask API → Rules engine: apply, check the end
  • Flask API → Three AI levels: AI reply
  • Three AI levels → Rules engine: tries moves
  • Hard AI weights → Three AI levels
RoleProject manager
What I doWeekly progress meetings with our supervisor and reviews with the client; market research; success criteria and a go-to-market strategy for Steam, itch.io and a browser demo.
RoleTechnical lead
What I doThe 10 × 8 engine and its tests; rule re-checks; the Hard AI and its move rules; my Unity game as the client; in-game feedback while playing; the research and plan for removing Flask.

Why a 12-number AI replaced a 635,712-number network

The earlier version came with a deep-learning agent, but it could not be reused. It was sized for an 8 × 8 board and never saw painted squares, which decide who can move. Only 7 of its 4,032 outputs even meant the same move on the 10 × 8 board.

The earlier model against the Hard AI
Earlier 8 × 8 networkHard AI
Learned numbers635,71212
Inputs65 (the 10 × 8 game needs 161)12 features that work on any board size
Outputs4,032 move scores (10 × 8 needs 6,320)one score for a position
Sees painted squaresNoYes
Runs inside Unity without extra softwareNo: needs a model runtimeYes: twelve multiplications

Network shapes decoded from the saved checkpoint; the 10 × 8 requirements derived from the rebuilt environment.

So I trained a different kind of AI. The Hard level tries every legal move on a copy of the game, measures the result with 12 features and scores it with 12 learned weights. The weights were fitted in seconds by ridge regression over positions from about 500 self-play games. It is not a neural network, and it does not learn while you play.

Explanatory diagramDerived from src/rl_agent.py, src/features.py and the accepted AI behaviour baseline (2 September 2026)

How the Hard AI chooses a move

How the Hard AI chooses a moveOne-move lookahead: try every legal move, score what each leaves behind, and play the best. The two tie rules only choose between moves the scores already rate (almost) equal.The position on theboardList the legal movesabout 52 on an average turn,always in the same orderPlay each move on acopy of the gameMeasure 12 featuresroyals alive, on goal andadvanced; elements alive;share painted; can each sidestill winScore = Σ weight ×feature12 multiplications per moveThe 12 weightsfitted by ridge regression onpositions from about 500self-play gamesKeep the moves with thebest scoreBreak exact tiesprefer a move that does notundo my last one, then theearliest listedBlock shuttlinginstead of a second reversalin a row, take anon-reversing move within0.0007Play the chosen move
How the Hard AI chooses a moveOne-move lookahead: try every legal move, score what each leaves behind, and play the best. The two tie rules only choose between moves the scores already rate (almost) equal.The position on theboardList the legal movesabout 52 on an averageturn, always in the sameorderPlay each move on acopy of the gameMeasure 12 featuresroyals alive, on goal andadvanced; elementsalive; share painted; caneach side still winScore = Σ weight ×feature12 multiplications permoveThe 12 weightsfitted by ridge regressionon positions from about500 self-play gamesKeep the moves withthe best scoreBreak exact tiesprefer a move that doesnot undo my last one,then the earliest listedBlock shuttlinginstead of a secondreversal in a row, take anon-reversing movewithin 0.0007Play the chosen move
  • Implemented
  • Stored data
One-move lookahead: try every legal move, score what each leaves behind, and play the best. The two tie rules only choose between moves the scores already rate (almost) equal.

How the Hard AI chooses a move

100%
How the Hard AI chooses a moveThe position on theboardList the legal movesabout 52 on an average turn,always in the same orderPlay each move on acopy of the gameMeasure 12 featuresroyals alive, on goal andadvanced; elements alive;share painted; can each sidestill winScore = Σ weight ×feature12 multiplications per moveThe 12 weightsfitted by ridge regression onpositions from about 500self-play gamesKeep the moves with thebest scoreBreak exact tiesprefer a move that does notundo my last one, then theearliest listedBlock shuttlinginstead of a second reversalin a row, take anon-reversing move within0.0007Play the chosen move
Text version of this diagram

Parts

  • The position on the board
  • List the legal moves — about 52 on an average turn, always in the same order
  • Play each move on a copy of the game
  • Measure 12 features — royals alive, on goal and advanced; elements alive; share painted; can each side still win
  • Score = Σ weight × feature — 12 multiplications per move
  • The 12 weights — fitted by ridge regression on positions from about 500 self-play games
  • Keep the moves with the best score
  • Break exact ties — prefer a move that does not undo my last one, then the earliest listed
  • Block shuttling — instead of a second reversal in a row, take a non-reversing move within 0.0007
  • Play the chosen move

Connections

  • The position on the board → List the legal moves
  • List the legal moves → Play each move on a copy of the game
  • Play each move on a copy of the game → Measure 12 features
  • Measure 12 features → Score = Σ weight × feature
  • The 12 weights → Score = Σ weight × feature: weights
  • Score = Σ weight × feature → Keep the moves with the best score
  • Keep the moves with the best score → Break exact ties
  • Break exact ties → Block shuttling
  • Block shuttling → Play the chosen move
legal moves in a recorded position
60
royal progress after the chosen move
0.25 → 0.33
the weight on royal progress
× 0.367
the chosen move's exact winning margin
+0.0306

A worked decision from the project's AI explainer, computed with the live engine. Every move the Hard AI makes can be explained the same way.

Chart · recorded resultsRead from trained_value_agent_10x8.json, the shipped Hard AI (fitted on the 10 × 8 engine)

Everything the Hard AI knows: twelve weights

The twelve weights of the Hard AI, from most positive to most negative
FeatureWeight
I can still win (read as a pair)my_can_win+0.804
How far my royals have advancedmy_royal_progress+0.367
My royals on their goal rankmy_royals_on_goal+0.311
Opponent's royals still on the board (read as a pair)opp_royals_alive+0.210
Opponent's elements still on the boardopp_elements_alive+0.171
Constantbias−0.008
Share of the board paintedpainted_fraction−0.051
My elements still on the boardmy_elements_alive−0.154
Opponent's royals on their goal rankopp_royals_on_goal−0.161
My royals still on the board (read as a pair)my_royals_alive−0.251
How far the opponent's royals have advancedopp_royal_progress−0.416
Opponent can still win (read as a pair)opp_can_win−0.766
  • the AI seeks it
  • the AI avoids it
  • read as a pair
Positive weights are things the AI seeks, negative ones things it avoids. The strongest lesson is a mirror image: stay able to win, and stop the opponent being able to win. Painted territory barely registers, which is the AI's known weak spot.

Rows marked ◆ should be read in pairs. “Can still win” and “royals still on the board” measure almost the same thing, so their credit is split between them: together they come to +0.553 for me and −0.556 for the opponent. The two element weights point the “wrong” way; the project flags them as possibly noise from a 500-game training run rather than explaining them away.

What testing established

80 automated tests cover the rules, checks against the original game's behaviour, the three AI levels and the server. The Hard AI was benchmarked over 40 games against each easier level, alternating colours: 85% against Easy (34 wins, 0 losses, 6 draws) and 57.5% against Normal (23–7–10).

Watching it play showed a problem the win rates hid: it kept moving the same piece back and forth. Many moves scored exactly the same, and the first-listed one kept flipping. Where a reversal did win, it won by a sliver, because going back over its own trail painted nothing new. Two rules fixed it: when scores tie exactly, prefer a move that does not undo the last one; and rather than make a second reversal in a row, take a non-reversing move scoring within 0.0007 of it.

The back-and-forth fix, measured
MeasureBeforeAfter
Hard AI moves that undo its previous move53.7%4.3%
Longest back-and-forth run19 moves2 moves
Win rate against Easy86.7% (60 games)85.0% (40 games)
Win rate against Normal55.0% (60 games)57.5% (40 games)
Illegal moves00

The win-rate confidence intervals overlap, so the honest claim is “not measurably weaker”, not “better”. No shuttle longer than two moves appeared in 25 live games through the Flask API.

Choosing how to remove Flask

The game still needs a Python server running on the player's machine. That rules out the plan to sell OrOgins on Steam and itch.io and to publish a self-contained browser demo. I compared six ways of removing it, scored them against eleven weighted criteria, and stress-tested the result.

The key finding was that Flask serves the rules, not just the AI: four of its five endpoints are rule operations. A model runtime such as Unity's Sentis cannot generate a legal move or paint a tile, so every option that removes Flask needs the rules in C#.

Chart · recorded resultsWeighted decision matrix from the no-Flask decision research (August 2026): 11 criteria, scores normalised to 0–100

Six ways to remove Flask, scored

Weighted scores for the six options, out of 100
OptionScore out of 100
HybridAcceptedC# rules and AI ship; Python kept as reference90.4
Port rules and AI to C#same shipped build, Python dropped82.0
C# search AIa stronger AI, but needs the port first74.2
Keep hosted Flaskworks, but never offline51.0
Bundle Python locallyPC only; no browser or mobile45.4
Sentis / ONNXruns models, not rules24.8
The hybrid led in all six scenario weightings tested and in all 20,000 randomised ones. Scored pessimistically on effort, risk and maintenance, the plain port edges ahead, 82.0 to 81.2: the hybrid is only worth it if the Python-versus-C# checks actually run.

The hybrid was accepted on 2 September 2026. The rules and all three AI levels move into C# inside Unity, with the 12 weights embedded as plain numbers. Python stops shipping but stays the authority: it records reference games, the C# version must match them at every step, and it remains the training environment and a possible future online server.

Explanatory diagramDerived from ARCHITECTURE_DECISION.md (ADR 001, accepted 2 September 2026) and the architecture options study (August 2026)

The accepted target: rules and AI inside Unity, Python kept as the reference

The accepted target: rules and AI inside Unity, Python kept as the referenceWhat players install contains no server and no model runtime. Python stops shipping but stays the authority: it records reference games, and the C# version must match them move for move. Dashed blue parts are planned, not built.WHAT PLAYERS INSTALL · UNITY AND C#NOT SHIPPEDUnity viewboard, pieces and menu,unchangedBackend interfacethe same five operations astoday's APILocal backendmethod calls instead of HTTPOrOgins.Core in C#rules, move generator and thethree AI levels; the 12 weightsembeddedHTTP backendtoday's client, kept fordevelopment and future onlineplayPython engine and Flaskthe golden reference, the trainingenvironment, a possible onlineserverGolden fixturesseeded games: board, orderedmoves, features, previous andchosen moveParity check in CIC# must match Python at everystepshipsdevelopment togglechecks
The accepted target: rules and AI inside Unity, Python kept as the referenceWhat players install contains no server and no model runtime. Python stops shipping but stays the authority: it records reference games, and the C# version must match them move for move. Dashed blue parts are planned, not built.SHIPPED · C#NOT SHIPPEDUnity viewboard, pieces and menu,unchangedBackend interfacethe same five operationsas today's APILocal backendmethod calls instead ofHTTPOrOgins.Core in C#rules, move generatorand the three AI levels;the 12 weightsembeddedHTTP backendtoday's client, kept fordevelopment and futureonline playPython engine andFlaskthe golden reference, thetraining environment, apossible online serverGolden fixturesseeded games: board,ordered moves, features,previous and chosenmoveParity check in CIC# must match Python atevery step
  • Implemented
  • Stored data
  • Planned or proposed — not built
What players install contains no server and no model runtime. Python stops shipping but stays the authority: it records reference games, and the C# version must match them move for move. Dashed blue parts are planned, not built.

The 12 weights are embedded as plain numbers in C#. Converting them to ONNX for Unity's model runtime was ruled out: it would wrap twelve multiplications in a model file and could change which of two near-tied moves wins.

The accepted target: rules and AI inside Unity, Python kept as the reference

100%
The accepted target: rules and AI inside Unity, Python kept as the referenceWHAT PLAYERS INSTALL · UNITY AND C#NOT SHIPPEDUnity viewboard, pieces and menu,unchangedBackend interfacethe same five operations astoday's APILocal backendmethod calls instead of HTTPOrOgins.Core in C#rules, move generator and thethree AI levels; the 12 weightsembeddedHTTP backendtoday's client, kept fordevelopment and future onlineplayPython engine and Flaskthe golden reference, the trainingenvironment, a possible onlineserverGolden fixturesseeded games: board, orderedmoves, features, previous andchosen moveParity check in CIC# must match Python at everystepshipsdevelopment togglechecks
Text version of this diagram

Parts

  • What players install · Unity and C#
    • Unity view — board, pieces and menu, unchanged
    • Backend interface — the same five operations as today's API; planned, not built
    • Local backend — method calls instead of HTTP; planned, not built
    • OrOgins.Core in C# — rules, move generator and the three AI levels; the 12 weights embedded; planned, not built
  • Not shipped
    • HTTP backend — today's client, kept for development and future online play; planned, not built
    • Python engine and Flask — the golden reference, the training environment, a possible online server
    • Golden fixtures — seeded games: board, ordered moves, features, previous and chosen move; planned, not built
    • Parity check in CI — C# must match Python at every step; planned, not built

Connections

  • Unity view → Backend interface
  • Backend interface → Local backend: ships (planned, not built)
  • Local backend → OrOgins.Core in C# (planned, not built)
  • Backend interface → HTTP backend: development toggle (planned, not built)
  • HTTP backend → Python engine and Flask
  • Python engine and Flask → Golden fixtures: records (planned, not built)
  • Golden fixtures → Parity check in CI (planned, not built)
  • Parity check in CI → OrOgins.Core in C#: checks (planned, not built)

Where it stands, and what I learned

Today the game is playable through the Python server. The shared codebase has moved to the 10 × 8 board and the C# port is under way. Progress is reviewed weekly with our supervisor, and an experienced reviewer will play the first version before wider testing.

  1. Freeze the data contract, the order legal moves are listed in, and the Hard AI's move rules.
  2. Record reference games from the Python engine.
  3. Port the rules, then the game flow, then the three AI levels to C#.
  4. Check the C# version against Python, move for move, automatically.
  5. Build offline PC and browser versions; mobile after that.
What Year 2 showedThe game and the model never met.
What the rebuild does about itUnity became a client of the Python engine from the start, and one interface will let the engine move into C#.
What Year 2 showedRules lived in two places and neither was complete.
What the rebuild does about itOne engine owns every rule, backed by 80 tests; Unity holds none.
What Year 2 showedThe agent was rewarded for legal moves, not good ones.
What the rebuild does about itA small evaluator trained on game outcomes, where every move can be explained.
What Year 2 showedTests checked the rules, never the policy.
What the rebuild does about itThe AI is benchmarked head-to-head, and its behaviour is frozen so a port can be checked against it.

Contact

I’m looking for graduate and early-career roles in AI/ML and software engineering.