hpr4746 :: Monotone
Lee does some Audio and Textual Analysis in Python
Hosted by Lee on Monday, 2026-10-12 is flagged as Clean and is released under a CC-BY-SA license.
python, audio, community, podcast.
(Be the first).
Listen in ogg,
opus,
or mp3 format. Play now:
Duration: 00:11:42
Download the transcription and
subtitles.
general.
Python and its libraries
1. librosa — loading and resampling audio
def load_audio(path): y, sr = librosa.load(path, sr=SR, mono=True) return y, sr
Every episode, regardless of its original sample rate or channel count, gets loaded down to a single mono stream at 16kHz (SR = 16000). That's one line of librosa, resampling, and folding stereo to mono. It made every downstream tool (VAD, pitch tracking, MFCC extraction) operate on a consistent, predictable input regardless of what an MP3 encoder or podcast host originally used.
2. parselmouth (Praat bindings) — pitch and voice-quality tracking
def pitch_and_hnr(y, sr): snd = parselmouth.Sound(y, sampling_frequency=sr) pitch = snd.to_pitch(time_step=PITCH_TIME_STEP, pitch_floor=PITCH_FLOOR, pitch_ceiling=PITCH_CEILING) f0 = pitch.selected_array["frequency"] times = pitch.xs() try: harmonicity = snd.to_harmonicity_cc(time_step=PITCH_TIME_STEP) hnr = np.array([harmonicity.get_value(t) for t in times]) hnr = np.nan_to_num(hnr, nan=-999.0) exceptException: hnr = np.full_like(f0, 999.0) # fail open if HNR extraction errors return f0, hnr, times
parselmouth is a Python wrapper around Praat, the standard tool in phonetics research. This is the heart of the whole vocal-monotony measurement: to_pitch() gives fundamental frequency (F0, roughly "pitch") every 20ms, and to_harmonicity_cc() gives harmonics-to-noise ratio (HNR) at the same timestamps — a signal-quality score used later to throw out pitch estimates taken from noisy or badly-compressed audio rather than let bad audio masquerade as "flat" pitch.
3. webrtcvad — deciding which frames are actually speech
def vad_speech_mask(y, sr, aggressiveness=2):
vad = webrtcvad.Vad(aggressiveness)
pcm16 = (np.clip(y, -1.0, 1.0) *32767).astype(np.int16)
n_frames =len(pcm16) // VAD_FRAME_LEN
mask = np.zeros(n_frames, dtype=bool)
for i inrange(n_frames):
chunk = pcm16[i * VAD_FRAME_LEN:(i +1) * VAD_FRAME_LEN]
raw = struct.pack("%dh"%len(chunk), *chunk)
try:
mask[i] = vad.is_speech(raw, sr)
exceptException:
mask[i] =False
return mask, times
webrtcvad is the voice-activity detector out of Google's WebRTC project — a frame-by-frame classifier ("is this 30ms chunk speech or not"). It needs raw 16-bit PCM bytes, not floats, hence the np.clip and struct.pack conversion.
4. librosa.feature — telling music apart from speech
flatness = librosa.feature.spectral_flatness(y=y, hop_length=hop)[0] chroma = librosa.feature.chroma_stft(y=y, sr=sr, hop_length=hop) onset_env = librosa.onset.onset_strength(y=y, sr=sr, hop_length=hop) ... chroma_diff = np.mean(np.linalg.norm(np.diff(chroma_win, axis=1), axis=0)) chroma_stability =1.0/ (1.0+ chroma_diff) ... scores[w] =0.4* chroma_stability +0.35* periodicity +0.25*min(flat_mean *3, 1.0)
Podcast intros/outros are often music, and music has a very different pitch profile than speech — including it would bias the "monotone" measurement. Rather than train a classifier, three librosa features were combined with hand-tuned weights: spectral flatness (music tends to have more tonal structure, less noise-like spectrum), chroma stability (sustained harmony vs.
constantly-changing speech formants), and onset periodicity (a regular beat vs. speech's irregular rhythm).
5. librosa.feature.mfcc + sklearn.cluster.KMeans + sklearn.metrics.silhouette_score — an unsupervised speaker-count attempt
mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=13, hop_length=hop) ... chunk_vecs.append(np.concatenate([seg.mean(axis=1), seg.std(axis=1)])) ... X = np.array(chunk_vecs) X = (X - X.mean(axis=0)) / (X.std(axis=0) +1e-9) best_k, best_score =1, -1.0 for k inrange(2, min(MAX_SPEAKERS, len(X) -1) +1): labels = KMeans(n_clusters=k, n_init=10, random_state=0).fit_predict(X) iflen(set(labels)) <2: continue score = silhouette_score(X, labels) if score > best_score: best_k, best_score = k, score estimated = best_k if best_score >= SILHOUETTE_MIN else1
MFCCs (Mel-frequency cepstral coefficients) summarize the shape of a short audio window's spectrum — a rough "voice fingerprint" widely used in speech tasks. The idea: chop each episode into 2-second chunks, turn each into a fingerprint vector, standardize (z-score) the features, then use KMeans to see if the chunks naturally split into 2+ clusters (different voices) rather than 1 (a single voice throughout). silhouette_score scores how well-separated the clusters are, and picks the best k — or falls back to "1 speaker" if no clustering clears a minimum quality bar. This part of the investigation did not work, as it picked up noise as another host.
6. scipy.stats.mannwhitneyu + a hand-written Cliff's delta — comparing two groups without assuming a normal distribution
def cliffs_delta(a, b): a, b = np.asarray(a), np.asarray(b) gt =sum((x > y) for x in a for y in b) lt =sum((x < y) for x in a for y in b) return (gt - lt) / (len(a) *len(b)) def compare_group(df, metric, label): hpr = df[df.corpus =="hpr"][metric].dropna() cmp= df[df.corpus =="comparison"][metric].dropna() u, p = stats.mannwhitneyu(hpr, cmp, alternative="two-sided") delta = cliffs_delta(hpr.values, cmp.values) iflen(hpr) *len(cmp) <200000elseNone
scipy.stats.mannwhitneyu is a statistical test: it compares two groups of numbers (HPR vs. comparison podcasts, per metric) without assuming they're normally distributed — appropriate for skewed, real-world data like pitch-range measurements. scipy doesn't ship an effect-size measure alongside it, so Cliff's delta — "of all possible (HPR episode, comparison episode) pairs, how much more often is HPR's value the smaller one?" This pairing (a significance test plus an effect-size number) was used throughout because a low p-value alone doesn't say whether the difference is actually big enough to matter.
7. pandas — grouping, stratifying, and cross-tabulating results
print(df.groupby("corpus")["retained_fraction"].describe())
...
print(pd.crosstab(df["corpus"], df["register"]))
Once data/results.jsonl (one JSON object per analyzed episode) is loaded into a pandas.DataFrame, two one-liners answer two different questions that would otherwise need custom code: .groupby(...).describe() gives a distribution summary (count, mean, std, quartiles) of audio-quality coverage per corpus, in one call, and pd.crosstab() cross-tabulates HPR vs. comparison against male-typical/female-typical register — a quick sanity check that the two collections of podcasts had a similar gender-register mix.
8. matplotlib — box plots for the headline comparison
fig, ax = plt.subplots(figsize=(6, 4)) data = [df[(df.corpus == c) & df[metric].notna()][metric].values for c in ["hpr", "comparison"]] ax.boxplot(data, tick_labels=["HPR", "Comparison"], showmeans=True)
A box plot per pitch-range metric, HPR next to comparison podcasts, mean marked as well as the median.
9. spacy word vectors — measuring topic drift, not just word choice
nlp = get_nlp() # spacy.load("en_core_web_md")
...
chunk_docs = [nlp(c) for c in chunks]
chunk_vecs = []
for d in chunk_docs:
v = d.vector
norm = np.linalg.norm(v)
if norm >1e-6:
chunk_vecs.append(v / norm)
...
adj_sims = [float(np.dot(V[i], V[i +1])) for i inrange(n -1)]
Each ~130-word chunk of a transcript is embedded as a single vector using spaCy's medium English model (en_core_web_md, which has built-in word vectors), then L2-normalized so a dot product between two vectors is their cosine similarity. Comparing consecutive chunks' similarity gives a "does the topic barely move between one chunk and the next" score.
10. lexicalrichness — vocabulary diversity that doesn't cheat on transcript length
lr = LexicalRichness(text) try: mtld = lr.mtld(threshold=0.72) exceptException: mtld =None
A naive vocabulary-diversity score (unique words ÷ total words) gets worse the longer a text runs, because of how the metric is defined — not because the speaker is actually repeating themselves more.
MTLD (Measure of Textual Lexical Diversity) avoids that by sweeping through the text and counting how many words it takes before diversity drops below a threshold, then averaging — a metric designed specifically to stay comparable between a 10-minute and a 40-minute episode.
11. vaderSentiment — a lexicon-based sentiment score
vader = get_vader() # SentimentIntensityAnalyzer() sentiments = [vader.polarity_scores(c)["compound"] for c in chunks] ... "sentiment_std": float(np.std(sentiments)) iflen(sentiments) >1elseNone,
VADER scores each chunk's sentiment from -1 to +1 using a hand-built lexicon and rule set (not a neural network) — chosen for being fast, deterministic, and cheap enough to run per-chunk across hundreds of episodes. The value then used isn't the average sentiment — it's the standard deviation of sentiment across an episode's chunks: does the emotional
tone ever move, or does it stay flat throughout?
12. zlib (standard library) — compression ratio as a gauge for entropy
def compression_ratio(text):
raw = text.encode("utf-8")
iflen(raw) <200:
returnNone
compressed = zlib.compress(raw, level=9)
returnlen(compressed) /len(raw)
No third-party NLP library involved at all — just the standard-library zlib module. A gzip-style compressor finds and exploits repeated patterns; if it can squash a transcript down a lot, that transcript was relatively repetitive and predictable, and if it can barely compress it, the wording was more varied. This is roughly tells you "verbal monotony."
13. re (standard library) — turning "does HPR favour solo hosts" into a checkable heuristic
MULTI_PRESENTER_PATTERNS = [
r"\bco-?host(?:s|ed|ing)?\b",
r"\bguest\b",
r"\binterview(?:s|ed|ing)?\b",
r"\btalks? with\b",
r"\bjoins? me\b",
...
]
CORRESPONDENT_LINK_RE = re.compile(r'correspondents/(\d+)\.html')
def multi_presenter_score(summary, notes):
text = (summary or"") +" "+ (notes or"")
hits = [p for p in MULTI_PRESENTER_PATTERNS if re.search(p, text, re.IGNORECASE)]
summary_links =set(CORRESPONDENT_LINK_RE.findall(summary or""))
return {
"flagged_multi_presenter": bool(hits) orlen(summary_links) >1,
}
Rather than the acoustic speaker-clustering attempt (item 5, which turned out to be unreliable), a more trustworthy stat came from plain text: HPR's own episode metadata (pulled from a local knowledge-base repo) has host-written summaries, and a short list of regex patterns ("guest", "interview", "joins me", etc.) plus counting linked correspondent profiles was enough to flag likely multi-presenter episodes.
14. multiprocessing.Pool — running the same analysis over hundreds of files without waiting hundreds of times over
work = [(r, local_paths[r["id"]]) for r in batch] with mp.Pool(args.workers) as pool, open(args.out, "a") as out_f: for result in pool.imap_unordered(_worker, work): out_f.write(json.dumps(result) +"\n") out_f.flush()
Every batch pipeline (acoustic pitch analysis, speaker-count pass) used the same shape: Pool(workers).imap_unordered(...) to run the per-episode function across multiple CPU cores at once, streaming results out to a JSON-lines file as each one finishes rather than waiting for the whole batch. imap_unordered was chosen over map so a single slow episode couldn't hold up writing out the results that had already finished.
15. subprocess + a resumable batch loop — working within a disk-space limit
def load_done_ids(results_path): ifnot os.path.exists(results_path): returnset() done =set() withopen(results_path) as f: for line in f: ... done.add(json.loads(line)["id"]) return done ... rows = [r for r in rows if r["id"] notin done] ... for batch_start inrange(0, len(rows), args.batch_size): batch = rows[batch_start:batch_start + args.batch_size] local_paths = fetch_batch(batch, args.scratch) # scp from the remote host ... for p in local_paths.values(): os.remove(p) # delete before next batch
All the podcasts (~220GB) didn't fit on local disk (~54GB free). The fix was a loop: scp (via subprocess) one batch of audio in, analyze it, delete it, move to the next batch. load_done_ids() makes the whole pipeline resumable by reading back what's already in the output file, so a crash or a deliberate stop partway doesn't lose completed stuff.
Links
Ah yes, the riveting world of Hacker Public Radio: a quaint little corner where tech wizards ✨️ divulge their magical secrets ✨️ to each other in the most riveting monotone known to man. Who knew endless "community" podcasts could be so... utterly riveting? ?️?️ Apparently, everyone is just dying to contribute to this cacophony once a year—what a privilege! ?
https://hackerpublicradio.org/
#HackerPublicRadio #techCommunity #podcasting #MonotoneMagic #rivetingContent #HackerNews #ngated