Samik Hafeez
All work

Year 1 team project · AI Project Design and Development

From podcasts and webpages to sentiment data

A three-person project that turned web articles, PDFs and podcasts into one text dataset for sentiment analysis. I defined the requirements, shaped the pipeline as Software Architect and built the speech-to-text route.

Context
AI Project Design and Development · academic scenario
My role
Business Analyst · Software Architect · Scrum Master, Sprint 1
Team
3 students
Dates
February – May 2024
Status
Completed coursework prototype

Stack

  • Python
  • Deepgram
  • BeautifulSoup
  • Selenium
  • PyMuPDF
  • NLTK
  • VADER
  • scikit-learn
  • pandas

The short version

Each point is expanded, with diagrams, in the sections below.

01Problem and constraints
An academic scenario asked whether a video-analytics company should enter the metaverse. Public discussion was spread across webpages, PDFs, forums and podcasts, each in a different format.
02My responsibility
Business Analyst and Software Architect: requirements, aims and the shape of the pipeline. Scrum Master for Sprint 1, when the data had to be gathered.
03The design
Convert every source into the same CSV text format, so one pipeline could clean, label and model all of it.
04What I implemented
The podcast route: Deepgram transcription after Google's recogniser failed my test, transcript-to-CSV conversion and cleaning of the speech text.
05The team’s part
Teammates scraped the web and PDF sources, ran the repository and built the four classifiers.
06What was delivered
A merged, labelled dataset; four classifiers compared on it; a final presentation.
07What testing established
My transcription test chose Deepgram over Google. The classifiers were scored against lexicon labels, with only a small human check.
08Learned, and unfinished
Agreeing data formats is engineering work. Pseudo-labels, a leaky vectoriser and a filter that dropped “not” limited what the scores meant.
On this page

A business question with messy data

The brief was an academic scenario: would it be worth a video-analytics company expanding into virtual stores? To answer it, we had to find out what people were actually saying, and they were saying it in articles, reports, comment threads and podcasts.

Each source brought its own problem. Dynamic websites defeated simple scraping, PDFs carried page headers into the text, forums blocked automated access, and Twitter refused API access after a week. Podcasts held some of the richest discussion but had no text at all.

My role, and the team's

Who did what, from the group blog, sprint plan and contribution records
AreaMy partTeammates
AnalysisBusiness Analyst: aims, objectives, functional and non-functional requirements; early company research.Project plan, test plan, risk table and quality plan.
ArchitectureSoftware Architect: one shared text format for every source.Repository set-up and development / QA ownership.
Data collectionPodcast transcription with Deepgram; transcript-to-CSV conversion.Web scraping (BeautifulSoup, Selenium) and PDF extraction (PyMuPDF).
PreparationCleaning of the speech transcripts.Cleaning of written sources; VADER labelling; merging.
Modelling—Bag-of-words and TF-IDF features; four classifiers.
DeliveryScrum Master for Sprint 1 (data gathering and requirements).Scrum Masters for Sprint 2, Sprint 3 and closure.

One pipeline for three kinds of source

The architecture decision was simple but it shaped everything: whatever the source, it had to arrive as rows of text in the same CSV format. Written sources went through scraping or PDF extraction; spoken sources went through my transcription route. After that, one set of cleaning, labelling and modelling steps applied to all of it.

Explanatory diagramDerived from my individual report, the group blog, the test plan and the final presentation

Three kinds of source, one shared text pipeline

Three kinds of source, one shared text pipelineEvery source became rows of text in one CSV format before cleaning and modelling. The tinted route is mine; the dashed blue items were planned but never happened.SOURCESSHARED PREPARATION (TEAM)MY SPEECH-TO-TEXT ROUTELABELLING AND MODELLING (TEAM)Web articlesrequests + BeautifulSoup;Selenium for dynamic pagesPDF reportsPyMuPDF text extractionForum commentsLinkedIn scraped; Quoraand Reddit copied by handTwitter / Xplanned; API access refusedPodcasts8 episodes as MP3Deepgram transcriptionpunctuation, speakers,paragraphsTranscripts → CSVone text file per episode,mergedClean speech textlower-case, letters only,drop “speaker” tags,stopwordsClean written textempty rows, duplicates,page headersShared CSV datasetone text column per rowVADER labelslexicon pseudo-labels;TextBlob comparedBoW and TF-IDFvocabulary raised from 100to 1,000 termsFour classifiersnaive Bayes, SVM, logisticregression, random forestCompare + human checkseveral split ratios; ~30hand-labelled commentsDeploymentoutside the projectaudiotext rowsspeech rows
Three kinds of source, one shared text pipelineEvery source became rows of text in one CSV format before cleaning and modelling. The tinted route is mine; the dashed blue items were planned but never happened.SOURCESPREPARATION (TEAM)MY SPEECH ROUTEMODELLING (TEAM)Web articlesrequests +BeautifulSoup; Seleniumfor dynamic pagesPDF reportsPyMuPDF text extractionForum commentsLinkedIn scraped; Quoraand Reddit copied byhandTwitter / Xplanned; API accessrefusedPodcasts8 episodes as MP3Deepgramtranscriptionpunctuation, speakers,paragraphsTranscripts → CSVone text file per episode,mergedClean speech textlower-case, letters only,drop “speaker” tags,stopwordsClean written textempty rows, duplicates,page headersShared CSV datasetone text column per rowVADER labelslexicon pseudo-labels;TextBlob comparedBoW and TF-IDFvocabulary raised from100 to 1,000 termsFour classifiersnaive Bayes, SVM, logisticregression, random forestCompare + humancheckseveral split ratios; ~30hand-labelled commentsDeploymentoutside the project
  • Implemented
  • My contribution
  • Stored data
  • Incomplete, or a gap found in testing
  • Planned or proposed — not built
Every source became rows of text in one CSV format before cleaning and modelling. The tinted route is mine; the dashed blue items were planned but never happened.

Web and PDF text came through scraping and extraction; forum comments partly by hand when platforms blocked access. Podcasts had to be transcribed first. From the shared dataset onwards, one pipeline labelled and modelled everything.

Three kinds of source, one shared text pipeline

100%
Three kinds of source, one shared text pipelineSOURCESSHARED PREPARATION (TEAM)MY SPEECH-TO-TEXT ROUTELABELLING AND MODELLING (TEAM)Web articlesrequests + BeautifulSoup;Selenium for dynamic pagesPDF reportsPyMuPDF text extractionForum commentsLinkedIn scraped; Quoraand Reddit copied by handTwitter / Xplanned; API access refusedPodcasts8 episodes as MP3Deepgram transcriptionpunctuation, speakers,paragraphsTranscripts → CSVone text file per episode,mergedClean speech textlower-case, letters only,drop “speaker” tags,stopwordsClean written textempty rows, duplicates,page headersShared CSV datasetone text column per rowVADER labelslexicon pseudo-labels;TextBlob comparedBoW and TF-IDFvocabulary raised from 100to 1,000 termsFour classifiersnaive Bayes, SVM, logisticregression, random forestCompare + human checkseveral split ratios; ~30hand-labelled commentsDeploymentoutside the projectaudiotext rowsspeech rows
Text version of this diagram

Parts

  • Sources
    • Web articles — requests + BeautifulSoup; Selenium for dynamic pages; teammates
    • PDF reports — PyMuPDF text extraction; teammates
    • Forum comments — LinkedIn scraped; Quora and Reddit copied by hand; teammates
    • Twitter / X — planned; API access refused; planned, not built
    • Podcasts — 8 episodes as MP3
  • Shared preparation (team)
    • Clean written text — empty rows, duplicates, page headers
    • Shared CSV dataset — one text column per row
  • My speech-to-text route
    • Deepgram transcription — punctuation, speakers, paragraphs; my contribution
    • Transcripts → CSV — one text file per episode, merged; my contribution
    • Clean speech text — lower-case, letters only, drop “speaker” tags, stopwords; my contribution
  • Labelling and modelling (team)
    • VADER labels — lexicon pseudo-labels; TextBlob compared
    • BoW and TF-IDF — vocabulary raised from 100 to 1,000 terms; teammate: modelling
    • Four classifiers — naive Bayes, SVM, logistic regression, random forest; teammate: modelling
    • Compare + human check — several split ratios; ~30 hand-labelled comments; incomplete or failed in testing
  • Deployment — outside the project; planned, not built

Connections

  • Web articles → Clean written text
  • PDF reports → Clean written text
  • Forum comments → Clean written text
  • Twitter / X → Clean written text (planned, not built)
  • Podcasts → Deepgram transcription: audio
  • Deepgram transcription → Transcripts → CSV
  • Transcripts → CSV → Clean speech text
  • Clean written text → Shared CSV dataset: text rows
  • Clean speech text → Shared CSV dataset: speech rows
  • Shared CSV dataset → VADER labels
  • VADER labels → BoW and TF-IDF
  • BoW and TF-IDF → Four classifiers
  • Four classifiers → Compare + human check
  • Compare + human check → Deployment (planned, not built)

My speech-to-text route

I started with Google's speech recogniser, and my test plan caught the problem straight away: it returned only fragments of each episode. I switched to Deepgram's pre-recorded API with punctuation, speaker separation and paragraphs enabled, which produced full, consistent transcripts. Each episode was written to a text file, the files were merged and converted into the shared CSV format, and I cleaned the speech text for analysis.

Explanatory diagramDerived from my test plan and the transcription code shown in my individual report

How I got usable text out of the podcasts

How I got usable text out of the podcastsMy speech-to-text route, including the failed first attempt. The test plan's first case failed with Google's recogniser; the second passed with Deepgram.Requirementspoken discussion must join thewritten sources as textTest 1 · Google speechrecognitionFull transcript?Only partial segmentstest failedTest 2 · Deepgrampre-recorded APIAPI key from .env; one request perepisodeFull, consistent transcript?Paragraph transcript → .txtMerge → CSVjoins the shared datasetKnown limitsneeds an API key and a stableconnection; no word-error ratemeasurednotry another serviceyes
How I got usable text out of the podcastsMy speech-to-text route, including the failed first attempt. The test plan's first case failed with Google's recogniser; the second passed with Deepgram.Requirementspoken discussion mustjoin the written sources astextTest 1 · Google speechrecognitionFull transcript?Only partial segmentstest failedTest 2 · Deepgrampre-recorded APIAPI key from .env; onerequest per episodeFull, consistenttranscript?Paragraph transcript →.txtMerge → CSVjoins the shared datasetKnown limitsneeds an API key and astable connection; noword-error rate measuredyes
  • Implemented
  • My contribution
  • Decision
  • Incomplete, or a gap found in testing
My speech-to-text route, including the failed first attempt. The test plan's first case failed with Google's recogniser; the second passed with Deepgram.

How I got usable text out of the podcasts

100%
How I got usable text out of the podcastsRequirementspoken discussion must join thewritten sources as textTest 1 · Google speechrecognitionFull transcript?Only partial segmentstest failedTest 2 · Deepgrampre-recorded APIAPI key from .env; one request perepisodeFull, consistent transcript?Paragraph transcript → .txtMerge → CSVjoins the shared datasetKnown limitsneeds an API key and a stableconnection; no word-error ratemeasurednotry another serviceyes
Text version of this diagram

Parts

  • Requirement — spoken discussion must join the written sources as text; my contribution
  • Test 1 · Google speech recognition — my contribution
  • Full transcript?
  • Only partial segments — test failed; incomplete or failed in testing
  • Test 2 · Deepgram pre-recorded API — API key from .env; one request per episode; my contribution
  • Full, consistent transcript?
  • Paragraph transcript → .txt — my contribution
  • Merge → CSV — joins the shared dataset; my contribution
  • Known limits — needs an API key and a stable connection; no word-error rate measured; incomplete or failed in testing

Connections

  • Requirement → Test 1 · Google speech recognition
  • Test 1 · Google speech recognition → Full transcript?
  • Full transcript? → Only partial segments: no
  • Only partial segments → Test 2 · Deepgram pre-recorded API: try another service
  • Test 2 · Deepgram pre-recorded API → Full, consistent transcript?
  • Full, consistent transcript? → Paragraph transcript → .txt: yes
  • Paragraph transcript → .txt → Merge → CSV
  • Test 2 · Deepgram pre-recorded API → Known limits (incomplete or failed in testing)

CRISP-DM, applied to real activities

We structured the project with CRISP-DM. In practice most of the effort sat in data understanding and preparation: platforms limited access, so collection routes had to change as we went, and every source needed its own cleaning. Deployment was deliberately outside the coursework.

Explanatory diagramDerived from my individual report and the group blog

CRISP-DM, mapped to what the team actually did

CRISP-DM, mapped to what the team actually didCRISP-DMacademic scenario: should a video-analyticscompany enter the metaverse?01Business understandingScenario, aims and functional /non-functional requirements02Data understandingCollect web, PDF, forum andpodcast text; change routeswhen platforms block access03Data preparationClean each source, VADERlabels, merge into one CSV04ModellingBoW / TF-IDF with fourclassifiers (teammate-led)05EvaluationCompare scores; smallhand-labelled check06DeploymentDeliberately left out of thecoursework
  1. Business understandingScenario, aims and functional / non-functional requirements
  2. Data understandingCollect web, PDF, forum and podcast text; change routes when platforms block access
  3. Data preparationClean each source, VADER labels, merge into one CSV
  4. ModellingBoW / TF-IDF with four classifiers (teammate-led)
  5. EvaluationCompare scores; small hand-labelled check
  6. DeploymentDeliberately left out of the coursework
  • Stage
  • My contribution
  • Incomplete, or a gap found in testing
  • Not done
CRISP-DM stages mapped to the project's actual activities. Tinted stages are where my analysis and data work sat; deployment was deliberately left out.

CRISP-DM, mapped to what the team actually did

100%
CRISP-DM, mapped to what the team actually didCRISP-DMacademic scenario: should a video-analyticscompany enter the metaverse?01Business understandingScenario, aims and functional /non-functional requirements02Data understandingCollect web, PDF, forum andpodcast text; change routeswhen platforms block access03Data preparationClean each source, VADERlabels, merge into one CSV04ModellingBoW / TF-IDF with fourclassifiers (teammate-led)05EvaluationCompare scores; smallhand-labelled check06DeploymentDeliberately left out of thecoursework
Text version of this diagram
  1. Business understanding: Scenario, aims and functional / non-functional requirements (my contribution)
  2. Data understanding: Collect web, PDF, forum and podcast text; change routes when platforms block access (my contribution)
  3. Data preparation: Clean each source, VADER labels, merge into one CSV
  4. Modelling: BoW / TF-IDF with four classifiers (teammate-led)
  5. Evaluation: Compare scores; small hand-labelled check
  6. Deployment: Deliberately left out of the coursework (not done)

What the evaluation can and cannot show

The team compared naive Bayes, SVM, logistic regression and random forest on bag-of-words and TF-IDF features, over several train/test ratios. Every score, though, measures agreement with VADER's lexicon labels rather than with human judgement, and the project's own records disagree about how the “balanced” dataset was produced. I therefore do not quote a headline accuracy.

Three weaknesses I would fix first
  • Labels: the only human check was a small set of about thirty comments judged by the team. A fair test needs an independently labelled holdout.
  • Leakage: the model notebook fits the text vectoriser before splitting, so the vocabulary sees the test text. The vectoriser belongs inside the training pipeline.
  • Cleaning: dropping words of three letters or fewer also removes “not”, which can flip the sentiment of a sentence. Negation should be kept and cleaning checked on examples.

What I took from it

This was the project where I learned that getting data into a usable, agreed shape is engineering work in its own right, and that a transcript can be well formatted and still wrong. It also gave me an early example of evaluating the whole pipeline rather than choosing a classifier in isolation.

Contact

I’m looking for graduate and early-career roles in AI/ML and software engineering.