Final-year client project for Bradford Council · 2025–26
Citizen Support AI Agent
A prototype that answers council-service questions from public Council pages and guides residents through bounded tasks, such as finding a bin day or booking an appointment, in one conversation. I was the team's Technical Lead.
- Context
- Final-year client project for Bradford Council, University of Bradford
- My role
- Technical Lead · first-semester Product Owner · Scrum Master, Sprints 2 and 4
- Team
- 3 students
- Dates
- Oct 2025 – Apr 2026
- Status
- Prototype demonstrated to the client, April 2026
Stack
- C# / ASP.NET Core
- Python / FastAPI
- LangChain + FAISS
- sentence-transformers
- OpenAI gpt-4o-mini
- Playwright
- xUnit / Moq
- Docker Compose
- AWS EC2

The short version
Each point is expanded, with diagrams, in the sections below.
- 01Problem and constraints
- Residents describe problems in their own words. Some need information; others need a lookup tied to their address. The brief was open, and we had no access to internal Council data.
- 02My responsibility
- Technical Lead. I proposed the approach, drafted the code structure, built the first MVP and several workflows, and led two sprints.
- 03The design
- C# decides every route: safety checks first, then step-by-step workflows, then retrieval. A Python service only searches Council text and words the answer.
- 04What I implemented
- Page selection and crawling for the knowledge base; the MVP; school finder, appointments and council-tax calculator; the bin lookup with a teammate; deployment work.
- 05The team’s part
- One teammate led research, the knowledge-base structure and the poster. Another integrated LangChain and automated the per-service tests.
- 06What was delivered
- A multi-service prototype in Docker on AWS EC2, demonstrated to the client and at the final exhibition.
- 07What testing established
- An archived 160-question routing run matched 108. More than half of the misses came from a test session stuck mid-booking, not from wording.
- 08Learned, and unfinished
- Design around the data you really have and test conversations with clean state. Live integrations, a privacy notice and better interruption handling remain.
On this page
An open brief and no internal data
The Council wanted a chatbot that could point residents to the right information. With three of us, six months and no access to Council systems, the first engineering problem was deciding what we could build and still stand behind.
I proposed using the information the Council already publishes as a retrieval knowledge base, and putting several services behind one conversation: questions answered from those pages, and bounded tasks handled step by step. I saw this as the point where I stepped up as a technical lead: turning a signposting brief into an applied-AI system we could build within the constraint.
- ConstraintNo internal Council records or accounts.
- Design responseGround answers in public service pages; keep bookings and school data as clearly labelled prototype data.
- ConstraintResidents use everyday words, not service names.
- Design responseDetect the service from keywords, then retrieve within that service before asking the model to write anything.
- ConstraintSome answers depend on an address.
- Design responseA deterministic workflow asks for a postcode and reads the Council's public bin-date form.
- ConstraintSome messages signal a crisis.
- Design responseA fixed safety reply runs before any other check or model call.
My responsibility, and the team's
My main focus was the technical build of a large, multi-service system. I also carried delivery roles: Product Owner in the first semester, turning broad client goals into a prioritised backlog, and Scrum Master for Sprints 2 and 4.
| Area | My part | Teammates |
|---|---|---|
| Direction and architecture | Proposed public-page retrieval with one chat for several services; drafted the code structure, class diagram and final design map. | Research methodology; sequence and use-case diagrams. |
| Knowledge base | Selected the Council service pages in scope and built the crawling automation. | Knowledge-base structure; Council links gathered during research. |
| Conversation and workflows | First MVP; school finder, appointment booking, council-tax calculator; official links in answers; bin lookup (shared). | Initial LLM agent; bin lookup (shared); LangChain integration. |
| Testing | Refactoring and code review of the final codebase. | Automated per-service evaluation scripts and test documentation. |
| Delivery | Docker, service networking, Nginx/HTTPS and AWS EC2 deployment work; a reproducible README. | GitHub repository, demonstration plan and submission. |
| Project management | Project plan, quality plan and ROAM risk analysis; Product Owner (first semester); Scrum Master (Sprints 2 and 4). | Charter, backlog, status reports; Scrum Master in the other sprints. |
How the system fits together
Every chat message goes to one ASP.NET Core endpoint. The C# orchestrator runs the safety and input checks, advances any workflow in progress, and only then decides whether retrieval is needed. The Python side does two narrow jobs: turning text into vectors, and searching the indexed Council pages before a model writes the answer.
How the Citizen Support AI prototype is put together
- Implemented
- My contribution
- External service or data
- Stored data
Three containers ran on one EC2 instance: the ASP.NET Core application, a small embedding service and the LangChain service. The Council's own website appears twice: as the source of the crawled knowledge base, and as the live bin-date form the Playwright workflow reads.
Text version of this diagram
Parts
- Resident's browser
- Chat interface — chat, topic chips, workflow panels
- Offline knowledge build
- Targeted crawler — 46 selected pages across 8 services; my contribution; page selection and crawling automation (my account)
- Clean & chunk — relevance filter, headings, metadata
- AWS EC2 · Docker Compose (demonstration)
- Nginx — HTTPS reverse proxy (in reports)
- ASP.NET Core · C#
- API controllers — /api/chat plus postcode, school, booking, form, voice
- ChatOrchestrator — guards, then workflows, then retrieval; my contribution; I drafted the code structure and built the first MVP; the team extended it
- Conversation state — per session, 30-min expiry, masked details
- FAQ retrieval gate — curated FAQ vectors, top 4, threshold
- Specialist workflows
- School finder — stored catalogue; my contribution
- Appointments — mock slots, in memory; my contribution
- Bin-collection lookup — Playwright via /api/postcode; my contribution; co-developed with a teammate
- Council-tax calculator — guidance arithmetic; my contribution
- Housing & forms — decision tree; forms not submitted
- Nearby services — distance to stored locations
- Python services · FastAPI
- FAISS index — chunks of council pages
- LangChain service — /agent: rule checks, search, rerank, wording; a teammate led the LangChain integration
- Embedding service — all-MiniLM-L6-v2
- bradford.gov.uk — public service pages
- OpenAI API — gpt-4o-mini
- Council bin-date form — public web form
- postcodes.io — postcode coordinates
Connections
- Chat interface → Nginx: HTTPS
- Nginx → API controllers
- API controllers → ChatOrchestrator: each message
- ChatOrchestrator → Conversation state: read / write
- ChatOrchestrator → workflows: one step per message
- ChatOrchestrator → FAQ retrieval gate: answerable?
- FAQ retrieval gate → Embedding service: /embed
- ChatOrchestrator → LangChain service: /agent + service hint
- LangChain service → FAISS index: search
- LangChain service → Embedding service: query vector
- LangChain service → OpenAI API: word the answer
- ChatOrchestrator → OpenAI API: fallback if empty
- Targeted crawler → bradford.gov.uk: crawl
- Targeted crawler → Clean & chunk: pages
- Clean & chunk → FAISS index: build index
- Bin-collection lookup → Council bin-date form: headless browser
- Nearby services → postcodes.io: coordinates
Keep workflows and retrieval apart
A question about council-tax discounts can be answered from published pages; a collection date needs a postcode and a live lookup. Routing between the two kept each path small enough to test.
Instead of One retrieve-and-generate prompt for every message.
Let code decide, and the model only word the answer
The route, the workflow steps and the official link are chosen in code. The model receives one or two retrieved chunks and an instruction to use only the facts they contain.
Instead of Letting an agent choose tools and routes on its own.
Put safeguarding before everything
Crisis wording is matched on the normalised text before any other check, so the reply is fixed and cannot be reworded or mixed with an earlier topic. Phrase matching is not exhaustive, and the case study treats it that way.
Two kinds of request
The clearest way to see the design is to follow two messages. An informational question passes the checks and goes to retrieval. A bin-day request never reaches a language model: it becomes a small state machine spread over several messages, with the browser and a headless browser doing the lookup.
An informational question and a service request take different routes
- Implemented
- Decision
- Incomplete, or a gap found in testing
- A · Informational: “What council-tax discounts are there?”
- B · Service request: “What day is my bin collected?” (two messages)
Informational questions and service requests share the same entry point but part ways early. Crisis wording always wins; workflows run before retrieval; retrieval only calls the Python service when the service is known or the match is strong.
The weak point is interruption. Cancel words reliably leave a booking, but a new question typed mid-booking can still be treated as the next booking answer: the evaluation below shows the same effect.
Text version of this diagram
Parts
- New message — normalised; service keywords detected
- Crisis wording?
- Fixed safety reply — NHS 111, Samaritans; stops here
- Small talk, vague or “something else”?
- Short reply or clarifying question — “something else” also clears pending steps
- Workflow pending or triggered? — bin day, booking, form, school, housing, tax calculator
- 1 · Ask for a postcode — pending step saved; no model
- 2 · Postcode → address → dates — UI signal, /api/postcode, Playwright
- Council form fails or changes — error shown as text; incomplete or failed in testing
- Cancel, answer or new question?
- Leave the flow — cancel words
- Next workflow step — deterministic C#; state saved
- New question mid-booking — meant to exit; can still be read as a booking answer; incomplete or failed in testing
- Embed + FAQ top 4 — within the detected service
- Confident match?
- Ask which service
- LangChain service — FAISS search, rerank, gpt-4o-mini wording
- OpenAI fallback — if the answer is empty
- Service unreachable — unhandled; the request fails; incomplete or failed in testing
- Reply — service label, official link, suggestions
Connections
- New message → Crisis wording?
- Crisis wording? → Fixed safety reply: yes
- Crisis wording? → Small talk, vague or “something else”?: no
- Small talk, vague or “something else”? → Short reply or clarifying question: yes
- Small talk, vague or “something else”? → Workflow pending or triggered?: no
- Workflow pending or triggered? → 1 · Ask for a postcode: bin day asked [B · Service request: “What day is my bin collected?” (two messages)]
- Workflow pending or triggered? → 2 · Postcode → address → dates: postcode after the prompt [B · Service request: “What day is my bin collected?” (two messages)]
- 2 · Postcode → address → dates → Council form fails or changes (incomplete or failed in testing)
- Workflow pending or triggered? → Cancel, answer or new question?: step in progress
- Cancel, answer or new question? → Leave the flow: cancel
- Cancel, answer or new question? → Next workflow step: answer
- Cancel, answer or new question? → New question mid-booking: new question (incomplete or failed in testing)
- Workflow pending or triggered? → Embed + FAQ top 4: no [A · Informational: “What council-tax discounts are there?”]
- Embed + FAQ top 4 → Confident match? [A · Informational: “What council-tax discounts are there?”]
- Confident match? → Ask which service: no
- Confident match? → LangChain service: yes [A · Informational: “What council-tax discounts are there?”]
- LangChain service → OpenAI fallback: empty
- LangChain service → Service unreachable: timeout (incomplete or failed in testing)
- LangChain service → Reply: answer + link [A · Informational: “What council-tax discounts are there?”]
- OpenAI fallback → Reply
- 2 · Postcode → address → dates → Reply: collection dates [B · Service request: “What day is my bin collected?” (two messages)]
- 1 · Ask for a postcode → Reply: prompt [B · Service request: “What day is my bin collected?” (two messages)]
- Next workflow step → Reply
A bin-collection request, call by call
- No language model is involved anywhere in this path.
- Resident to Chat UI“What day is my bin collected?”
- Chat UI to ChatOrchestratormessage + session id
- ChatOrchestrator to Conversation statepending step: awaiting postcode
- ChatOrchestrator to Chat UIreply“Please enter your postcode…”
- next message
- Resident to Chat UI“BD3 8PX”
- Chat UI to ChatOrchestratormessage + session id
- ChatOrchestrator to Conversation statemasked postcode; lookup started
- ChatOrchestrator to Chat UIreplyPOSTCODE_LOOKUP signal
- Chat UI to Postcode APIGET /api/postcode/search
- Postcode API to Council formheadless browser enters postcode
- Council form to Postcode APIreplyaddress options
- Postcode API to Chat UIreplyaddress buttons
- Resident to Chat UIchooses an address
- Chat UI to Postcode APIGET /api/postcode/bin-result
- Postcode API to Council formselect address, show dates
- Council form to Postcode APIreplycollection dates
- Postcode API to Conversation statestore result for follow-ups
- Postcode API to Chat UIreplynext general, recycling and garden dates
- If the form times out or changes, the error comes back as text: the lookup depends on a page the team did not control.
- Implemented
- My contribution
- External service or data
- Stored data
- Incomplete, or a gap found in testing
The chat endpoint never scrapes anything itself. It stores the pending step, then answers the postcode with a signal the browser turns into a separate lookup request, so the slow web lookup runs in its own request rather than inside the conversation handler.


Real screens from the prototype, April 2026.
From Council pages to an answer
Because we had no internal data, the knowledge base was only as good as the pages we chose. I selected the Council service pages that matched our scope, leaving unrelated topics out, and built the automation that crawled them. The team then filtered, chunked and indexed the text, and tuned how results were reranked.
From selected Council pages to a grounded answer
- Implemented
- My contribution
- External service or data
- Stored data
The C# gate decides whether a question is answerable and which service it belongs to. Only then does the Python service search the page index, rerank with explicit rules, and pass one or two chunks to the model with an instruction to use only those facts. The official link comes from code, not from the model.
Text version of this diagram
Parts
- Building the knowledge base
- Services in scope — 8 areas: council tax, bins, benefits, schools, planning, libraries, housing, contact
- Select service pages — 46 seed URLs; unrelated topics left out; my contribution; my account: I chose the pages
- Crawl the selection — Playwright; same domain; 2 links deep; expands accordions; my contribution; my account: I built the crawling automation
- Page text + service label — boilerplate stripped; JSON
- Relevance filter — keep in-scope services; drop navigation pages
- Heading-aware chunks — 1,000 chars (150 overlap); FAQ-style pages 500/60
- Embed + attach metadata — MiniLM vectors; service, topic, heading, URL
- FAISS index — saved to disk, loaded once
- At question time
- Question — plus the service detected in C#
- FAQ gate — cosine top 4; below threshold → clarify
- FAISS search — 36 candidates
- Rule-based rerank — boost service, topic, heading; penalise off-topic pages
- Context — top 1–2 chunks, up to 1,800 chars
- gpt-4o-mini writes the answer — “use only facts that appear below”
- Weak-answer check — hedging replies discarded
- Official page link — best-matching council URL
- Answer + link
- Curated FAQ path (C#)
- Curated FAQ entries — service, title, answer, next-step URL; knowledge-base structure led by a teammate
- Sentence chunks — up to 420 characters
- Cached FAQ vectors — re-embedded when the FAQ changes
Connections
- Services in scope → Select service pages
- Select service pages → Crawl the selection
- Crawl the selection → Page text + service label
- Page text + service label → Relevance filter
- Relevance filter → Heading-aware chunks
- Heading-aware chunks → Embed + attach metadata
- Embed + attach metadata → FAISS index
- Question → FAQ gate
- FAQ gate → FAISS search: confident
- FAISS index → FAISS search: similarity search
- FAISS search → Rule-based rerank
- Rule-based rerank → Context
- Context → gpt-4o-mini writes the answer
- gpt-4o-mini writes the answer → Weak-answer check
- Weak-answer check → Official page link
- Official page link → Answer + link
- Curated FAQ entries → Sentence chunks
- Sentence chunks → Cached FAQ vectors
- Cached FAQ vectors → FAQ gate: compare
For this data work we adapted CRISP-DM. It gave us a way to talk about source quality and evaluation, even though no model was trained: the “modelling” step was embedding, indexing and rule-based reranking.
CRISP-DM, adapted to the knowledge base rather than to model training
- Business understandingWhich resident questions and council services the prototype should cover
- Data understandingChoose the Council service pages in scope; leave unrelated topics out
- Data preparationCrawl the selection, strip boilerplate, filter and chunk
- ModellingEmbed chunks, build the index, tune rule-based reranking
- EvaluationRouting and response checks exposed weak services
- DeploymentContainers on AWS EC2 for the client demonstration
- Stage
- My contribution
- Incomplete, or a gap found in testing
Text version of this diagram
- Business understanding: Which resident questions and council services the prototype should cover
- Data understanding: Choose the Council service pages in scope; leave unrelated topics out (my contribution)
- Data preparation: Crawl the selection, strip boilerplate, filter and chunk (my contribution)
- Modelling: Embed chunks, build the index, tune rule-based reranking
- Evaluation: Routing and response checks exposed weak services
- Deployment: Containers on AWS EC2 for the client demonstration
How we delivered it
SDLC structured the engineering work from analysis to deployment, and Scrum gave it a rhythm of planning, reviews and retrospectives. Tracking combined Microsoft Project, status reports, a ROAM risk log and Jira. The timeline shows how the phases fell across the sprints, and what I delivered in each.
Six months of sprints, and where the SDLC phases fell
| Sprint | Team goal (sprint plan) | My part |
|---|---|---|
| Initiation21 Oct – 4 Jan | Governance: charter, backlog, ROAM risk analysis. | Project plan, quality plan and ROAM analysis; first-semester Product Owner. |
| Sprint 15 Jan – 5 Feb | Research and experiments; MVP direction set on bins and benefits. | — |
| Sprint 26 Feb – 19 Feb | Design: ethics, diagrams, costing, CDIO. | Scrum Master. Coordinated the design phase; class diagram, cost review and CDIO report. |
| Sprint 320 Feb – 23 Mar | First MVP, LLM API refinement, exhibition plan. | Built the first working MVP and refined it with the LLM API. |
| Sprint 424 Mar – 12 Apr | Bin scraping, LangChain, links, school / appointment / council-tax features, automated tests. | Scrum Master. School finder, appointment booking, council-tax calculator; bin lookup with a teammate; official links in answers. |
| Sprint 513 Apr – 21 Apr | README, refactoring, documentation, demonstration. | Refactoring, code review and a reproducible README. |
Initiation21 Oct – 4 Jan
Governance: charter, backlog, ROAM risk analysis.
My part Project plan, quality plan and ROAM analysis; first-semester Product Owner.
- SDLC phase Analysis & research
- My roles Technical Lead throughout · Product Owner (first semester)
Sprint 15 Jan – 5 Feb
Research and experiments; MVP direction set on bins and benefits.
- SDLC phase Analysis & research
- My roles Technical Lead throughout · Product Owner (first semester)
Sprint 26 Feb – 19 FebScrum Master: me
Design: ethics, diagrams, costing, CDIO.
My part Coordinated the design phase; class diagram, cost review and CDIO report.
- SDLC phase Design
- My roles Technical Lead throughout
Sprint 320 Feb – 23 Mar
First MVP, LLM API refinement, exhibition plan.
My part Built the first working MVP and refined it with the LLM API.
- SDLC phase Build: MVP, then features
- My roles Technical Lead throughout
Sprint 424 Mar – 12 AprScrum Master: me
Bin scraping, LangChain, links, school / appointment / council-tax features, automated tests.
My part School finder, appointment booking, council-tax calculator; bin lookup with a teammate; official links in answers.
- SDLC phase Build: MVP, then features · Test, deploy, refine
- My roles Technical Lead throughout
Sprint 513 Apr – 21 Apr
README, refactoring, documentation, demonstration.
My part Refactoring, code review and a reproducible README.
- SDLC phase Test, deploy, refine
- My roles Technical Lead throughout
- Sprint or phase
- My role or sprint
Text version of this diagram
- Initiation (21 Oct – 4 Jan): Governance: charter, backlog, ROAM risk analysis. My part: Project plan, quality plan and ROAM analysis; first-semester Product Owner.
- Sprint 1 (5 Jan – 5 Feb): Research and experiments; MVP direction set on bins and benefits.
- Sprint 2 (6 Feb – 19 Feb): Design: ethics, diagrams, costing, CDIO. I was Scrum Master. My part: Coordinated the design phase; class diagram, cost review and CDIO report.
- Sprint 3 (20 Feb – 23 Mar): First MVP, LLM API refinement, exhibition plan. My part: Built the first working MVP and refined it with the LLM API.
- Sprint 4 (24 Mar – 12 Apr): Bin scraping, LangChain, links, school / appointment / council-tax features, automated tests. I was Scrum Master. My part: School finder, appointment booking, council-tax calculator; bin lookup with a teammate; official links in answers.
- Sprint 5 (13 Apr – 21 Apr): README, refactoring, documentation, demonstration. My part: Refactoring, code review and a reproducible README.
- SDLC phase: Analysis & research (21 Oct – 5 Feb); Design (6 Feb – 19 Feb); Build: MVP, then features (20 Feb – 12 Apr); Test, deploy, refine (24 Mar – 21 Apr)
- My roles: Technical Lead throughout (21 Oct – 21 Apr); Product Owner (first semester) (21 Oct – 31 Jan)
Getting beyond localhost was its own piece of work. Locally the services talked to each other; in containers they failed on ports, missing environment variables, service names and HTTPS. I worked through those failures until the three services ran together on AWS EC2 for the demonstrations.
Leadership in practice: the proposal, as a STAR example
- Situation
- An open-ended brief for a signposting chatbot, and no access to internal Council data.
- Task
- Help define an architecture with enough applied-AI scope for a final-year project that we could still deliver.
- Action
- I proposed public Council pages as a retrieval knowledge base and one chat for several bounded service workflows, then drafted the code structure and built the first MVP.
- Result
- The team delivered a working multi-service prototype and demonstrated it to the client. The evaluation also showed where routing and conversation state still broke.
- Learning
- Separate what you can build from what you can verify, and design the evidence for a system's answers from the start.
What testing established
| Test | What it checks | What it cannot show |
|---|---|---|
| 160-question routing run | Whether each reply carries the expected service label: 20 questions for each of 8 services. | Answer quality. All questions share one session, so earlier turns can affect later ones. |
| Response checks | Routing, expected phrases, link and suggestion relevance. The submitted script has 21 cases; the report describes a 45-case run. | Factual correctness: phrase matches are proxies. |
| C# test suite | xUnit with Moq and stubbed HTTP: routing, memory, privacy masking, safeguarding, workflows, API contracts (12 areas). | Live model quality or production load. |
Archived routing run: 108 of 160 questions reached the expected service
| Service | Result | Passed |
|---|---|---|
| Council Tax | 0 passed, 20 appointment prompt, 0 other service, 0 error | 0/20 |
| Waste & Bins | 19 passed, 1 appointment prompt, 0 other service, 0 error | 19/20 |
| Benefits & Support | 17 passed, 0 appointment prompt, 3 other service, 0 error | 17/20 |
| Education | 17 passed, 0 appointment prompt, 3 other service, 0 error | 17/20 |
| Planning | 20 passed, 0 appointment prompt, 0 other service, 0 error | 20/20 |
| Libraries | 19 passed, 0 appointment prompt, 1 other service, 0 error | 19/20 |
| Housing | 10 passed, 0 appointment prompt, 9 other service, 1 error | 10/20 |
| Contact Us | 6 passed, 8 appointment prompt, 6 other service, 0 error | 6/20 |
- reached the expected service
- answered with the appointment prompt (stuck session)
- routed to another service
- request error
The script sends all 160 questions through one fixed session, in order. The very first Council Tax question already received the appointment prompt, so that session was left mid-booking before the run began; after question 152 (“How do I book an appointment?”) the last eight Contact Us questions were caught the same way.
So the weakest scores mix two different problems: genuine routing confusion (for example Housing questions labelled as Benefits) and conversation state that a new question did not clear. The project report quotes a higher 85% from a run whose raw results I could not find; this chart uses only the inspectable archive.
See how routing was scored
SESSION_ID = "automated-test-session"
def grade_service(actual: str, expected: str) -> str:
if (actual or "").strip().lower() == (expected or "").strip().lower():
return "PASS"
return "FAIL"A stronger run would give each single-turn question a fresh session, keep a separate multi-turn suite for follow-ups and interruptions, record the source revision, and publish a confusion matrix.
What I learned, and what remains
At the exhibition the client's main point was that a public version would need a clear privacy notice and a plan for handling real citizen data. That, and our own tests, shaped the next steps.
- What the prototype showedA new question during a booking could be read as a booking answer, and state leaked across test questions.
- What a next version needsReliable interruption handling, and routing tests with a fresh session per question.
- What the prototype showedHousing and Contact Us were the weakest areas, and safeguarding relies on phrase lists.
- What a next version needsImprove those routes first; test distress wording far more widely.
- What the prototype showedBookings, school data and council-tax amounts are prototype data.
- What a next version needsApproved live integrations, with authentication and data handling agreed with the Council.
- What the prototype showedPostcodes, emails and phone numbers are masked in logs, but that is not a privacy regime.
- What a next version needsThe privacy notice the client asked for, retention rules and an infrastructure review.