Year 1 team project · AI Project Design and Development
From podcasts and webpages to sentiment data
A three-person project that turned web articles, PDFs and podcasts into one text dataset for sentiment analysis. I defined the requirements, shaped the pipeline as Software Architect and built the speech-to-text route.
- Context
- AI Project Design and Development · academic scenario
- My role
- Business Analyst · Software Architect · Scrum Master, Sprint 1
- Team
- 3 students
- Dates
- February – May 2024
- Status
- Completed coursework prototype
Stack
- Python
- Deepgram
- BeautifulSoup
- Selenium
- PyMuPDF
- NLTK
- VADER
- scikit-learn
- pandas
The short version
Each point is expanded, with diagrams, in the sections below.
- 01Problem and constraints
- An academic scenario asked whether a video-analytics company should enter the metaverse. Public discussion was spread across webpages, PDFs, forums and podcasts, each in a different format.
- 02My responsibility
- Business Analyst and Software Architect: requirements, aims and the shape of the pipeline. Scrum Master for Sprint 1, when the data had to be gathered.
- 03The design
- Convert every source into the same CSV text format, so one pipeline could clean, label and model all of it.
- 04What I implemented
- The podcast route: Deepgram transcription after Google's recogniser failed my test, transcript-to-CSV conversion and cleaning of the speech text.
- 05The team’s part
- Teammates scraped the web and PDF sources, ran the repository and built the four classifiers.
- 06What was delivered
- A merged, labelled dataset; four classifiers compared on it; a final presentation.
- 07What testing established
- My transcription test chose Deepgram over Google. The classifiers were scored against lexicon labels, with only a small human check.
- 08Learned, and unfinished
- Agreeing data formats is engineering work. Pseudo-labels, a leaky vectoriser and a filter that dropped “not” limited what the scores meant.
On this page
A business question with messy data
The brief was an academic scenario: would it be worth a video-analytics company expanding into virtual stores? To answer it, we had to find out what people were actually saying, and they were saying it in articles, reports, comment threads and podcasts.
Each source brought its own problem. Dynamic websites defeated simple scraping, PDFs carried page headers into the text, forums blocked automated access, and Twitter refused API access after a week. Podcasts held some of the richest discussion but had no text at all.
My role, and the team's
| Area | My part | Teammates |
|---|---|---|
| Analysis | Business Analyst: aims, objectives, functional and non-functional requirements; early company research. | Project plan, test plan, risk table and quality plan. |
| Architecture | Software Architect: one shared text format for every source. | Repository set-up and development / QA ownership. |
| Data collection | Podcast transcription with Deepgram; transcript-to-CSV conversion. | Web scraping (BeautifulSoup, Selenium) and PDF extraction (PyMuPDF). |
| Preparation | Cleaning of the speech transcripts. | Cleaning of written sources; VADER labelling; merging. |
| Modelling | — | Bag-of-words and TF-IDF features; four classifiers. |
| Delivery | Scrum Master for Sprint 1 (data gathering and requirements). | Scrum Masters for Sprint 2, Sprint 3 and closure. |
One pipeline for three kinds of source
The architecture decision was simple but it shaped everything: whatever the source, it had to arrive as rows of text in the same CSV format. Written sources went through scraping or PDF extraction; spoken sources went through my transcription route. After that, one set of cleaning, labelling and modelling steps applied to all of it.
Three kinds of source, one shared text pipeline
- Implemented
- My contribution
- Stored data
- Incomplete, or a gap found in testing
- Planned or proposed — not built
Web and PDF text came through scraping and extraction; forum comments partly by hand when platforms blocked access. Podcasts had to be transcribed first. From the shared dataset onwards, one pipeline labelled and modelled everything.
Text version of this diagram
Parts
- Sources
- Web articles — requests + BeautifulSoup; Selenium for dynamic pages; teammates
- PDF reports — PyMuPDF text extraction; teammates
- Forum comments — LinkedIn scraped; Quora and Reddit copied by hand; teammates
- Twitter / X — planned; API access refused; planned, not built
- Podcasts — 8 episodes as MP3
- Shared preparation (team)
- Clean written text — empty rows, duplicates, page headers
- Shared CSV dataset — one text column per row
- My speech-to-text route
- Deepgram transcription — punctuation, speakers, paragraphs; my contribution
- Transcripts → CSV — one text file per episode, merged; my contribution
- Clean speech text — lower-case, letters only, drop “speaker” tags, stopwords; my contribution
- Labelling and modelling (team)
- VADER labels — lexicon pseudo-labels; TextBlob compared
- BoW and TF-IDF — vocabulary raised from 100 to 1,000 terms; teammate: modelling
- Four classifiers — naive Bayes, SVM, logistic regression, random forest; teammate: modelling
- Compare + human check — several split ratios; ~30 hand-labelled comments; incomplete or failed in testing
- Deployment — outside the project; planned, not built
Connections
- Web articles → Clean written text
- PDF reports → Clean written text
- Forum comments → Clean written text
- Twitter / X → Clean written text (planned, not built)
- Podcasts → Deepgram transcription: audio
- Deepgram transcription → Transcripts → CSV
- Transcripts → CSV → Clean speech text
- Clean written text → Shared CSV dataset: text rows
- Clean speech text → Shared CSV dataset: speech rows
- Shared CSV dataset → VADER labels
- VADER labels → BoW and TF-IDF
- BoW and TF-IDF → Four classifiers
- Four classifiers → Compare + human check
- Compare + human check → Deployment (planned, not built)
My speech-to-text route
I started with Google's speech recogniser, and my test plan caught the problem straight away: it returned only fragments of each episode. I switched to Deepgram's pre-recorded API with punctuation, speaker separation and paragraphs enabled, which produced full, consistent transcripts. Each episode was written to a text file, the files were merged and converted into the shared CSV format, and I cleaned the speech text for analysis.
How I got usable text out of the podcasts
- Implemented
- My contribution
- Decision
- Incomplete, or a gap found in testing
Text version of this diagram
Parts
- Requirement — spoken discussion must join the written sources as text; my contribution
- Test 1 · Google speech recognition — my contribution
- Full transcript?
- Only partial segments — test failed; incomplete or failed in testing
- Test 2 · Deepgram pre-recorded API — API key from .env; one request per episode; my contribution
- Full, consistent transcript?
- Paragraph transcript → .txt — my contribution
- Merge → CSV — joins the shared dataset; my contribution
- Known limits — needs an API key and a stable connection; no word-error rate measured; incomplete or failed in testing
Connections
- Requirement → Test 1 · Google speech recognition
- Test 1 · Google speech recognition → Full transcript?
- Full transcript? → Only partial segments: no
- Only partial segments → Test 2 · Deepgram pre-recorded API: try another service
- Test 2 · Deepgram pre-recorded API → Full, consistent transcript?
- Full, consistent transcript? → Paragraph transcript → .txt: yes
- Paragraph transcript → .txt → Merge → CSV
- Test 2 · Deepgram pre-recorded API → Known limits (incomplete or failed in testing)


Figures from the team's final presentation, May 2024.
CRISP-DM, applied to real activities
We structured the project with CRISP-DM. In practice most of the effort sat in data understanding and preparation: platforms limited access, so collection routes had to change as we went, and every source needed its own cleaning. Deployment was deliberately outside the coursework.
CRISP-DM, mapped to what the team actually did
- Business understandingScenario, aims and functional / non-functional requirements
- Data understandingCollect web, PDF, forum and podcast text; change routes when platforms block access
- Data preparationClean each source, VADER labels, merge into one CSV
- ModellingBoW / TF-IDF with four classifiers (teammate-led)
- EvaluationCompare scores; small hand-labelled check
- DeploymentDeliberately left out of the coursework
- Stage
- My contribution
- Incomplete, or a gap found in testing
- Not done
Text version of this diagram
- Business understanding: Scenario, aims and functional / non-functional requirements (my contribution)
- Data understanding: Collect web, PDF, forum and podcast text; change routes when platforms block access (my contribution)
- Data preparation: Clean each source, VADER labels, merge into one CSV
- Modelling: BoW / TF-IDF with four classifiers (teammate-led)
- Evaluation: Compare scores; small hand-labelled check
- Deployment: Deliberately left out of the coursework (not done)
What the evaluation can and cannot show
The team compared naive Bayes, SVM, logistic regression and random forest on bag-of-words and TF-IDF features, over several train/test ratios. Every score, though, measures agreement with VADER's lexicon labels rather than with human judgement, and the project's own records disagree about how the “balanced” dataset was produced. I therefore do not quote a headline accuracy.
Three weaknesses I would fix first
- Labels: the only human check was a small set of about thirty comments judged by the team. A fair test needs an independently labelled holdout.
- Leakage: the model notebook fits the text vectoriser before splitting, so the vocabulary sees the test text. The vectoriser belongs inside the training pipeline.
- Cleaning: dropping words of three letters or fewer also removes “not”, which can flip the sentiment of a sentence. Negation should be kept and cleaning checked on examples.
What I took from it
This was the project where I learned that getting data into a usable, agreed shape is engineering work in its own right, and that a transcript can be well formatted and still wrong. It also gave me an early example of evaluating the whole pipeline rather than choosing a classifier in isolation.