Year 2 team project, then a 2026 rebuild I lead
OrOgins: getting a game and its AI to work together
OrOgins is an elemental strategy board game whose digital version was created in 2015. The brief was to modernise its interface and let people play against an AI. In my Year 2 project the game and the AI were built separately and never connected. Since June 2026 I have led a rebuild as project manager and technical lead, and the game and its AI now work together. The next step, under way, is removing the Python server so the game runs offline.
- Context
- Year 2 university project, then a 2026 rebuild with client reviews
- My role
- Year 2: Project Manager, Unity development and testing · Rebuild: project manager and technical lead
- Team
- Year 2: 3 students · Rebuild: a small team with a supervisor, weekly progress meetings and client reviews
- Dates
- Nov 2024 – Apr 2025 · June 2026 – present
- Status
- Playable against three AI levels; still needs a local Python server; offline C# version in progress
Stack
- Unity 6
- C#
- Python
- NumPy
- Flask
- pytest
- Stable-Baselines3 (Year 2)
Scope of this pageThe rebuild is ongoing work with a client. Its figures come from the project's own test runs and recorded benchmarks (September 2026); the 80 tests were re-run, and passed, when this page was written.

The short version
Each point is expanded, with diagrams, in the sections below.
- 01Problem and constraints
- A 2015 digital board game needed a modern interface and an AI opponent. The AI had to follow the game's real rules and become part of the game, not a separate program.
- 02My responsibility
- Year 2: Project Manager and Scrum Master, Unity developer and tester. Rebuild: project manager and technical lead.
- 03The design
- Unity only draws the board and sends clicks; Python owns the rules and the AI. The accepted next design moves both into C# inside Unity and keeps Python as the reference and training environment.
- 04What I implemented
- Year 2: the Unity board, pieces, movement scripts and tests. Rebuild: the 10 × 8 engine, the 12-weight Hard AI and its move rules, my Unity game as the interface, and the research into removing Flask.
- 05The team’s part
- Year 2: a teammate built the Gym environment and trained PPO. The rebuild started from an earlier 8 × 8 student version, with a supervisor and client reviews along the way.
- 06What was delivered
- Year 2: a Unity game and a separate learning prototype. Rebuild: a playable game against Easy, Normal and Hard opponents, 80 passing tests, and an accepted plan for offline play.
- 07What testing established
- The Hard AI wins 85% of 40 games against Easy and 57.5% against Normal. A fix cut its back-and-forth moves from 53.7% to 4.3% with no measurable loss of strength.
- 08Learned, and unfinished
- Connect the halves first, build on the real rules before training anything, and prefer a model the team can explain. The game still needs Flask; the C# version is not finished.
On this page
The game, and the brief
OrOgins is a two-player elemental strategy game whose digital version was created in 2015. The brief, in my Year 2 project and again in the rebuild, was to modernise its interface and let people play it against an AI opponent.
Each side has two royals and eight element pieces: earth, water, fire and air, each beating the next in a cycle. Elements move like chess queens and paint the squares they cross. Royals can only travel forward or sideways along painted, empty squares, so the elements have to build their roads. A player wins by getting both royals to the rank in front of the opponent's home row. An element that crosses a square painted in the colour it beats wipes it and knocks off a royal standing there, and a player who loses a royal can no longer win.
The board both halves of the project had to agree on
Text version of this diagram
Ten columns by eight rows. Row 0 holds the Creationist pieces and row 7 the Evolutionist pieces, each in the order earth, water, fire, air, woman, man, air, fire, water, earth. Creationists aim for row 6 and Evolutionists for row 1. Element strength runs earth over water, water over fire, fire over air and air over earth.
· Phase 1 · 2024–25
Year 2: two halves that never met
I managed our three-person team and built the game in Unity; a teammate encoded the rules in a Python Gym environment and trained a PPO agent on it. The plan was to export the trained policy and run it inside Unity. The export was never written, nothing in Unity called the model wrapper, and when the module ended the game and the AI still could not talk to each other.
| Area | My part | Teammates |
|---|---|---|
| Management | Project Manager; Scrum Master for Sprints 1 and 3; project plan in Microsoft Project; risk assessment and quality plan. | Trello board, GitHub; Scrum Master for Sprint 2. |
| Requirements and design | Functional and non-functional requirements; chose Unity over Unreal; board and piece art. | Five-phase plan and player-experience questions. |
| Unity game | Unity set-up, board, piece prefabs, and the C# scripts for pieces, moves, capture and neutral squares. | Unity game logic assigned to a teammate; the element rules stayed unfinished. |
| Learning environment | Helped with the environment design and debugging. | Gym environment, state and action spaces, and the Pygame prototype rules. |
| Model | Model testing, debugging and monitoring training, with the teammate. | Chose and trained PPO with Stable-Baselines3. |
| Testing | Wrote, ran and documented the test and validation scripts; analysed the failures. | Model validation. |
Two systems that were meant to meet
- Implemented
- My contribution
- Stored data
- Incomplete, or a gap found in testing
- Planned or proposed — not built
The trained policy was saved in Stable-Baselines3's own format. Nothing exported it to ONNX, nothing in Unity called the Barracuda wrapper, and nothing mapped the Unity board to the environment's 100-value observation. That missing bridge is the project's main unfinished piece.
Text version of this diagram
Parts
- Unity game · C# (my build)
- Game.cs controller — creates 10 pieces a side; neutral squares fill the rest; my contribution
- Pieces + move plates — OrOrginsMan.cs, MovePlate.cs: click to move, capture; my contribution
- Neutral-square rule — experiment: fire pieces only; incomplete or failed in testing; my contribution
- Element, capture, victory and turn rules — not implemented in Unity; incomplete or failed in testing
- AIModelManager — Barracuda wrapper; never called; incomplete or failed in testing
- Planned integration (planned, not built)
- Export policy to ONNX — planned, not built
- Map Unity board → 100-value observation — planned, not built
- Python RL · teammate's environment and training
- OriginsEnv (Gym) — board rules, rewards, game-over checks; teammate
- PPO training — Stable-Baselines3 MlpPolicy, 100k steps configured; teammate
- ppo_origins.zip — saved SB3 model
- Pygame play loop — opponent plays random legal moves; teammate
- Test + validation scripts — 9 tests, 9 validators, 4 evaluators; my contribution; I wrote, ran and documented the model tests
Connections
- Game.cs controller → Pieces + move plates
- Pieces + move plates → Neutral-square rule
- Game.cs controller → Element, capture, victory and turn rules (incomplete or failed in testing)
- OriginsEnv (Gym) → PPO training: observations, rewards
- PPO training → ppo_origins.zip
- OriginsEnv (Gym) → Pygame play loop
- ppo_origins.zip → Export policy to ONNX (planned, not built)
- Export policy to ONNX → AIModelManager (planned, not built)
- Map Unity board → 100-value observation → AIModelManager (planned, not built)
- AIModelManager → Game.cs controller: chosen move (planned, not built)
- Test + validation scripts → OriginsEnv (Gym): exercise the rules
Year 2 in detail: what the agent learned from, what the tests showed, and how we ran it
Reading the environment code shows how narrow the agent's choice really was. It picks a square; the environment then moves that piece to the first legal destination it finds. Rewards say “legal move” or “illegal move” and ±100 at the end, so good play was never directly rewarded.
One step of the observation–action–reward loop, as coded
- Implemented
- Decision
- Incomplete, or a gap found in testing
Text version of this diagram
Parts
- PPO policy (MlpPolicy) — one policy plays both sides; teammate
- Action: a square, 0–79 — the policy never picks a destination
- Own piece on that square?
- Reward −1 — same player tries again
- Any legal move?
- Reward −1 — turn passes
- Take the first legal destination — fixed direction order
- Move, convert path squares, capture — reward +1
- Game over? — both arrived, a man or woman lost, or no moves
- Reward ±100 and reset — sign follows the mover, not the winner; incomplete or failed in testing
- Switch turn
- Observation: 100 integers — 80 squares, turn, 4 arrival flags, padding; bounds declared ±3 but pieces use ±6; incomplete or failed in testing
Connections
- PPO policy (MlpPolicy) → Action: a square, 0–79
- Action: a square, 0–79 → Own piece on that square?
- Own piece on that square? → Reward −1: no
- Own piece on that square? → Any legal move?: yes
- Any legal move? → Reward −1: no
- Any legal move? → Take the first legal destination: yes
- Take the first legal destination → Move, convert path squares, capture
- Move, convert path squares, capture → Game over?
- Game over? → Reward ±100 and reset: yes
- Game over? → Switch turn: no
- Switch turn → Observation: 100 integers
- Reward −1 → Observation: 100 integers
- Reward −1 → Observation: 100 integers
- Observation: 100 integers → PPO policy (MlpPolicy): next step
Without capture and victory rules in place, a win rate could not measure anything, so I tested what could be checked deterministically. The failures were the useful part: they pointed at missing rules rather than at the learning code.
Test and validation results recorded in the final reflection
- Board dimensions Passed10 columns × 8 rows
- Board initialisation MixedPieces and starting positionsTests passed; the stricter validator failed 1 of 4 checks.
- Observation encoding PassedLength and piece codes of the 100-value observation
- Element interactions MixedThe strength cycle, including fire against airTests passed; the air-against-earth validation crashed with an index error.
- Movement FailedPieces move in their allowed directions“Evolutionist man should move up”.
- Turn switching FailedThe turn passes after a move
- Destination timing MixedThe goal rows are recognisedStalemate after a maximum number of turns failed.
- AI decision checks MixedCapture and protective moves are availableChecks the rules, never the trained policy; strategic progress failed.
- Whole-game validation Mixed26 rule checks across the environment15 of 26 passed.
Several copies of the environment existed, so these results describe the build tested at the time, not a single final version.

We worked in Agile sprints with a Trello backlog, stand-ups and reviews; SDLC structured the game track and CRISP-DM the learning track. Sprint 1 ran long because Unity and the rules took far more time than planned, so by the time the environment and PPO arrived there was little room left to connect them.
Three sprints, two tracks, one unfinished connection
| Sprint | Team goal (sprint plan) | My part |
|---|---|---|
| Sprint 125 Nov – 16 Mar | Initialise the project, set up Unity, build the board and pieces. | Scrum Master. Requirements, project, risk and quality plans; Unity setup, board, pieces and C# scripts. |
| Sprint 217 Mar – 3 Apr | Custom Gym environment, PPO training and integration. | Model testing and debugging with the teammate who built the environment. |
| Sprint 34 Apr – 6 Apr | Completion, evaluation, validation and the presentation. | Scrum Master. Ran and documented the test and validation scripts; sprint review. |
Sprint 125 Nov – 16 MarScrum Master: me
Initialise the project, set up Unity, build the board and pieces.
My part Requirements, project, risk and quality plans; Unity setup, board, pieces and C# scripts.
- Game track (SDLC) Requirements, plans, Unity board and pieces
Sprint 217 Mar – 3 Apr
Custom Gym environment, PPO training and integration.
My part Model testing and debugging with the teammate who built the environment.
- Learning track (CRISP-DM) Gym environment + PPO (teammate)
- Integration Trained policy inside Unity: not reached
Sprint 34 Apr – 6 AprScrum Master: me
Completion, evaluation, validation and the presentation.
My part Ran and documented the test and validation scripts; sprint review.
- Integration Trained policy inside Unity: not reached
- Testing Tests + validation
- Sprint or work
- Mine
- Planned, not reached
Text version of this diagram
- Sprint 1 (25 Nov – 16 Mar): Initialise the project, set up Unity, build the board and pieces. I was Scrum Master. My part: Requirements, project, risk and quality plans; Unity setup, board, pieces and C# scripts.
- Sprint 2 (17 Mar – 3 Apr): Custom Gym environment, PPO training and integration. My part: Model testing and debugging with the teammate who built the environment.
- Sprint 3 (4 Apr – 6 Apr): Completion, evaluation, validation and the presentation. I was Scrum Master. My part: Ran and documented the test and validation scripts; sprint review.
- Game track (SDLC): Requirements, plans, Unity board and pieces (25 Nov – 16 Mar)
- Learning track (CRISP-DM): Gym environment + PPO (teammate) (17 Mar – 3 Apr)
- Integration: Trained policy inside Unity: not reached (17 Mar – 6 Apr)
- Testing: Tests + validation (4 Apr – 6 Apr)
CRISP-DM for the learning track
- Business understandingAn AI opponent for a two-player strategy game
- Data understandingRules, board state, actions and observations
- Data preparationRules coded into OriginsEnv; only legal moves allowed
- ModellingPPO with Stable-Baselines3 (teammate)
- EvaluationMy test and validation scripts; rule checks, not a win rate
- DeploymentPolicy inside the Unity game: not reached
- Stage
- My contribution
- Incomplete, or a gap found in testing
- Not reached
Text version of this diagram
- Business understanding: An AI opponent for a two-player strategy game
- Data understanding: Rules, board state, actions and observations
- Data preparation: Rules coded into OriginsEnv; only legal moves allowed
- Modelling: PPO with Stable-Baselines3 (teammate)
- Evaluation: My test and validation scripts; rule checks, not a win rate (my contribution)
- Deployment: Policy inside the Unity game: not reached (not done)
- Situation
- An unfamiliar engine, complex rules and two technologies developed side by side.
- Task
- Plan the work, build the Unity foundation and keep the team heading for a working game with an AI opponent.
- Action
- I set up the plans, built the board and scripts, and turned testing into a record of which rules worked and which did not.
- Result
- A Unity prototype and a separate trained-agent pipeline, but no working connection between them.
- Learning
- Plan smaller end-to-end milestones, agree one environment specification, and test the integration boundary first. It is exactly where the rebuild started.
· Phase 2 · June 2026 – present
The rebuild: one engine, and my game as its interface
In June 2026 I came back to OrOgins as project manager and technical lead of a rebuild, with a supervisor, weekly progress meetings and reviews with the client. The goal was the one Year 2 missed: a game you can actually play against the AI.
The starting point was an earlier 8 × 8 student version of the Python engine. The original game is played on 10 columns by 8 rows, so I reworked the engine to implement the original rules on the correct board: royals riding painted trails, captures by wiping tiles, and the win and can't-win conditions. Every rule became an automated test.
My Year 2 Unity game became the interface. It now holds no rules and no AI: it draws whatever state the Python engine sends back, and passes on the player's clicks. Players choose PvP or PvAI and Easy, Normal or Hard, and for the first time the game, in its proper interface, plays against a trained AI.
How the rebuilt game runs today
- Implemented
- Stored data
- Incomplete, or a gap found in testing
This split is what finally made the game and the AI work together, and it is also why the next step is possible: the Unity side already holds no rules, so the server can later be swapped for local C# code without touching what players see.
Text version of this diagram
Parts
- Unity 6 · my Year 2 game, now the interface
- Menu and status panel — PvP or PvAI; Easy, Normal or Hard; New Game
- Board, tiles and pieces — draws the game state it is sent, in my Year 2 art
- API client — no rules and no AI inside Unity
- Python process · must be started before the game
- Flask API — five endpoints: new_game, state, legal_moves, move, health
- Rules engine — 10 × 8, the original rules; 80 automated tests
- Three AI levels — Easy random · Normal greedy · Hard learned
- Hard AI weights — 12 numbers in a 677-byte file
- The remaining dependency — a second process has to be running: no offline play, no browser demo, no mobile; incomplete or failed in testing
Connections
- Menu and status panel → API client: new game
- Board, tiles and pieces → API client: clicks
- API client → Flask API: HTTP
- Flask API → Rules engine: apply, check the end
- Flask API → Three AI levels: AI reply
- Three AI levels → Rules engine: tries moves
- Hard AI weights → Three AI levels
- RoleProject manager
- What I doWeekly progress meetings with our supervisor and reviews with the client; market research; success criteria and a go-to-market strategy for Steam, itch.io and a browser demo.
- RoleTechnical lead
- What I doThe 10 × 8 engine and its tests; rule re-checks; the Hard AI and its move rules; my Unity game as the client; in-game feedback while playing; the research and plan for removing Flask.
Why a 12-number AI replaced a 635,712-number network
The earlier version came with a deep-learning agent, but it could not be reused. It was sized for an 8 × 8 board and never saw painted squares, which decide who can move. Only 7 of its 4,032 outputs even meant the same move on the 10 × 8 board.
| Earlier 8 × 8 network | Hard AI | |
|---|---|---|
| Learned numbers | 635,712 | 12 |
| Inputs | 65 (the 10 × 8 game needs 161) | 12 features that work on any board size |
| Outputs | 4,032 move scores (10 × 8 needs 6,320) | one score for a position |
| Sees painted squares | No | Yes |
| Runs inside Unity without extra software | No: needs a model runtime | Yes: twelve multiplications |
Network shapes decoded from the saved checkpoint; the 10 × 8 requirements derived from the rebuilt environment.
So I trained a different kind of AI. The Hard level tries every legal move on a copy of the game, measures the result with 12 features and scores it with 12 learned weights. The weights were fitted in seconds by ridge regression over positions from about 500 self-play games. It is not a neural network, and it does not learn while you play.
How the Hard AI chooses a move
- Implemented
- Stored data
Text version of this diagram
Parts
- The position on the board
- List the legal moves — about 52 on an average turn, always in the same order
- Play each move on a copy of the game
- Measure 12 features — royals alive, on goal and advanced; elements alive; share painted; can each side still win
- Score = Σ weight × feature — 12 multiplications per move
- The 12 weights — fitted by ridge regression on positions from about 500 self-play games
- Keep the moves with the best score
- Break exact ties — prefer a move that does not undo my last one, then the earliest listed
- Block shuttling — instead of a second reversal in a row, take a non-reversing move within 0.0007
- Play the chosen move
Connections
- The position on the board → List the legal moves
- List the legal moves → Play each move on a copy of the game
- Play each move on a copy of the game → Measure 12 features
- Measure 12 features → Score = Σ weight × feature
- The 12 weights → Score = Σ weight × feature: weights
- Score = Σ weight × feature → Keep the moves with the best score
- Keep the moves with the best score → Break exact ties
- Break exact ties → Block shuttling
- Block shuttling → Play the chosen move
- legal moves in a recorded position
- 60
- royal progress after the chosen move
- 0.25 → 0.33
- the weight on royal progress
- × 0.367
- the chosen move's exact winning margin
- +0.0306
A worked decision from the project's AI explainer, computed with the live engine. Every move the Hard AI makes can be explained the same way.
Everything the Hard AI knows: twelve weights
| Feature | Weight |
|---|---|
| I can still win (read as a pair)my_can_win | +0.804 |
| How far my royals have advancedmy_royal_progress | +0.367 |
| My royals on their goal rankmy_royals_on_goal | +0.311 |
| Opponent's royals still on the board (read as a pair)opp_royals_alive | +0.210 |
| Opponent's elements still on the boardopp_elements_alive | +0.171 |
| Constantbias | −0.008 |
| Share of the board paintedpainted_fraction | −0.051 |
| My elements still on the boardmy_elements_alive | −0.154 |
| Opponent's royals on their goal rankopp_royals_on_goal | −0.161 |
| My royals still on the board (read as a pair)my_royals_alive | −0.251 |
| How far the opponent's royals have advancedopp_royal_progress | −0.416 |
| Opponent can still win (read as a pair)opp_can_win | −0.766 |
- the AI seeks it
- the AI avoids it
- read as a pair
Rows marked ◆ should be read in pairs. “Can still win” and “royals still on the board” measure almost the same thing, so their credit is split between them: together they come to +0.553 for me and −0.556 for the opponent. The two element weights point the “wrong” way; the project flags them as possibly noise from a 500-game training run rather than explaining them away.
What testing established
80 automated tests cover the rules, checks against the original game's behaviour, the three AI levels and the server. The Hard AI was benchmarked over 40 games against each easier level, alternating colours: 85% against Easy (34 wins, 0 losses, 6 draws) and 57.5% against Normal (23–7–10).
Watching it play showed a problem the win rates hid: it kept moving the same piece back and forth. Many moves scored exactly the same, and the first-listed one kept flipping. Where a reversal did win, it won by a sliver, because going back over its own trail painted nothing new. Two rules fixed it: when scores tie exactly, prefer a move that does not undo the last one; and rather than make a second reversal in a row, take a non-reversing move scoring within 0.0007 of it.
| Measure | Before | After |
|---|---|---|
| Hard AI moves that undo its previous move | 53.7% | 4.3% |
| Longest back-and-forth run | 19 moves | 2 moves |
| Win rate against Easy | 86.7% (60 games) | 85.0% (40 games) |
| Win rate against Normal | 55.0% (60 games) | 57.5% (40 games) |
| Illegal moves | 0 | 0 |
The win-rate confidence intervals overlap, so the honest claim is “not measurably weaker”, not “better”. No shuttle longer than two moves appeared in 25 live games through the Flask API.
Choosing how to remove Flask
The game still needs a Python server running on the player's machine. That rules out the plan to sell OrOgins on Steam and itch.io and to publish a self-contained browser demo. I compared six ways of removing it, scored them against eleven weighted criteria, and stress-tested the result.
The key finding was that Flask serves the rules, not just the AI: four of its five endpoints are rule operations. A model runtime such as Unity's Sentis cannot generate a legal move or paint a tile, so every option that removes Flask needs the rules in C#.
Six ways to remove Flask, scored
| Option | Score out of 100 | |
|---|---|---|
| HybridAcceptedC# rules and AI ship; Python kept as reference | 90.4 | |
| Port rules and AI to C#same shipped build, Python dropped | 82.0 | |
| C# search AIa stronger AI, but needs the port first | 74.2 | |
| Keep hosted Flaskworks, but never offline | 51.0 | |
| Bundle Python locallyPC only; no browser or mobile | 45.4 | |
| Sentis / ONNXruns models, not rules | 24.8 |
The hybrid was accepted on 2 September 2026. The rules and all three AI levels move into C# inside Unity, with the 12 weights embedded as plain numbers. Python stops shipping but stays the authority: it records reference games, the C# version must match them at every step, and it remains the training environment and a possible future online server.
The accepted target: rules and AI inside Unity, Python kept as the reference
- Implemented
- Stored data
- Planned or proposed — not built
The 12 weights are embedded as plain numbers in C#. Converting them to ONNX for Unity's model runtime was ruled out: it would wrap twelve multiplications in a model file and could change which of two near-tied moves wins.
Text version of this diagram
Parts
- What players install · Unity and C#
- Unity view — board, pieces and menu, unchanged
- Backend interface — the same five operations as today's API; planned, not built
- Local backend — method calls instead of HTTP; planned, not built
- OrOgins.Core in C# — rules, move generator and the three AI levels; the 12 weights embedded; planned, not built
- Not shipped
- HTTP backend — today's client, kept for development and future online play; planned, not built
- Python engine and Flask — the golden reference, the training environment, a possible online server
- Golden fixtures — seeded games: board, ordered moves, features, previous and chosen move; planned, not built
- Parity check in CI — C# must match Python at every step; planned, not built
Connections
- Unity view → Backend interface
- Backend interface → Local backend: ships (planned, not built)
- Local backend → OrOgins.Core in C# (planned, not built)
- Backend interface → HTTP backend: development toggle (planned, not built)
- HTTP backend → Python engine and Flask
- Python engine and Flask → Golden fixtures: records (planned, not built)
- Golden fixtures → Parity check in CI (planned, not built)
- Parity check in CI → OrOgins.Core in C#: checks (planned, not built)
Where it stands, and what I learned
Today the game is playable through the Python server. The shared codebase has moved to the 10 × 8 board and the C# port is under way. Progress is reviewed weekly with our supervisor, and an experienced reviewer will play the first version before wider testing.
- Freeze the data contract, the order legal moves are listed in, and the Hard AI's move rules.
- Record reference games from the Python engine.
- Port the rules, then the game flow, then the three AI levels to C#.
- Check the C# version against Python, move for move, automatically.
- Build offline PC and browser versions; mobile after that.
- What Year 2 showedThe game and the model never met.
- What the rebuild does about itUnity became a client of the Python engine from the start, and one interface will let the engine move into C#.
- What Year 2 showedRules lived in two places and neither was complete.
- What the rebuild does about itOne engine owns every rule, backed by 80 tests; Unity holds none.
- What Year 2 showedThe agent was rewarded for legal moves, not good ones.
- What the rebuild does about itA small evaluator trained on game outcomes, where every move can be explained.
- What Year 2 showedTests checked the rules, never the policy.
- What the rebuild does about itThe AI is benchmarked head-to-head, and its behaviour is frozen so a port can be checked against it.