hpr4746 :: Monotone

Lee does some Audio and Textual Analysis in Python

Hosted by Lee on Monday, 2026-10-12 is flagged as Clean and is released under a CC-BY-SA license.
python, audio, community, podcast. (Be the first).

Listen in ogg, opus, or mp3 format. Play now:

Duration: 00:11:42
Download the transcription and subtitles.

general.

Python and its libraries

1. librosa — loading and resampling audio

def load_audio(path):
  y, sr = librosa.load(path, sr=SR, mono=True)
 return y, sr

Every episode, regardless of its original sample rate or channel count, gets loaded down to a single mono stream at 16kHz (SR = 16000). That's one line of librosa, resampling, and folding stereo to mono. It made every downstream tool (VAD, pitch tracking, MFCC extraction) operate on a consistent, predictable input regardless of what an MP3 encoder or podcast host originally used.

2. parselmouth (Praat bindings) — pitch and voice-quality tracking

def pitch_and_hnr(y, sr):
  snd = parselmouth.Sound(y, sampling_frequency=sr)
  pitch = snd.to_pitch(time_step=PITCH_TIME_STEP, pitch_floor=PITCH_FLOOR, pitch_ceiling=PITCH_CEILING)
  f0 = pitch.selected_array["frequency"]
  times = pitch.xs()
 try:
  harmonicity = snd.to_harmonicity_cc(time_step=PITCH_TIME_STEP)
  hnr = np.array([harmonicity.get_value(t) for t in times])
  hnr = np.nan_to_num(hnr, nan=-999.0)
 exceptException:
  hnr = np.full_like(f0, 999.0) # fail open if HNR extraction errors
 return f0, hnr, times

parselmouth is a Python wrapper around Praat, the standard tool in phonetics research. This is the heart of the whole vocal-monotony measurement: to_pitch() gives fundamental frequency (F0, roughly "pitch") every 20ms, and to_harmonicity_cc() gives harmonics-to-noise ratio (HNR) at the same timestamps — a signal-quality score used later to throw out pitch estimates taken from noisy or badly-compressed audio rather than let bad audio masquerade as "flat" pitch.

3. webrtcvad — deciding which frames are actually speech

def vad_speech_mask(y, sr, aggressiveness=2):
  vad = webrtcvad.Vad(aggressiveness)
  pcm16 = (np.clip(y, -1.0, 1.0) *32767).astype(np.int16)
  n_frames =len(pcm16) // VAD_FRAME_LEN
  mask = np.zeros(n_frames, dtype=bool)
 for i inrange(n_frames):
  chunk = pcm16[i * VAD_FRAME_LEN:(i +1) * VAD_FRAME_LEN]
  raw = struct.pack("%dh"%len(chunk), *chunk)
 try:
  mask[i] = vad.is_speech(raw, sr)
 exceptException:
  mask[i] =False
 return mask, times

webrtcvad is the voice-activity detector out of Google's WebRTC project — a frame-by-frame classifier ("is this 30ms chunk speech or not"). It needs raw 16-bit PCM bytes, not floats, hence the np.clip and struct.pack conversion.

4. librosa.feature — telling music apart from speech

flatness = librosa.feature.spectral_flatness(y=y, hop_length=hop)[0]
 chroma = librosa.feature.chroma_stft(y=y, sr=sr, hop_length=hop)
 onset_env = librosa.onset.onset_strength(y=y, sr=sr, hop_length=hop)
 ...
 chroma_diff = np.mean(np.linalg.norm(np.diff(chroma_win, axis=1), axis=0))
 chroma_stability =1.0/ (1.0+ chroma_diff)
 ...
 scores[w] =0.4* chroma_stability +0.35* periodicity +0.25*min(flat_mean *3, 1.0)

Podcast intros/outros are often music, and music has a very different pitch profile than speech — including it would bias the "monotone" measurement. Rather than train a classifier, three librosa features were combined with hand-tuned weights: spectral flatness (music tends to have more tonal structure, less noise-like spectrum), chroma stability (sustained harmony vs.

constantly-changing speech formants), and onset periodicity (a regular beat vs. speech's irregular rhythm).

5. librosa.feature.mfcc + sklearn.cluster.KMeans + sklearn.metrics.silhouette_score — an unsupervised speaker-count attempt

mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=13, hop_length=hop)
 ...
 chunk_vecs.append(np.concatenate([seg.mean(axis=1), seg.std(axis=1)]))
 ...
 X = np.array(chunk_vecs)
 X = (X - X.mean(axis=0)) / (X.std(axis=0) +1e-9)
 
 best_k, best_score =1, -1.0
 for k inrange(2, min(MAX_SPEAKERS, len(X) -1) +1):
  labels = KMeans(n_clusters=k, n_init=10, random_state=0).fit_predict(X)
 iflen(set(labels)) <2:
 continue
  score = silhouette_score(X, labels)
 if score > best_score:
  best_k, best_score = k, score
 
 estimated = best_k if best_score >= SILHOUETTE_MIN else1

MFCCs (Mel-frequency cepstral coefficients) summarize the shape of a short audio window's spectrum — a rough "voice fingerprint" widely used in speech tasks. The idea: chop each episode into 2-second chunks, turn each into a fingerprint vector, standardize (z-score) the features, then use KMeans to see if the chunks naturally split into 2+ clusters (different voices) rather than 1 (a single voice throughout). silhouette_score scores how well-separated the clusters are, and picks the best k — or falls back to "1 speaker" if no clustering clears a minimum quality bar. This part of the investigation did not work, as it picked up noise as another host.

6. scipy.stats.mannwhitneyu + a hand-written Cliff's delta — comparing two groups without assuming a normal distribution

def cliffs_delta(a, b):
  a, b = np.asarray(a), np.asarray(b)
  gt =sum((x > y) for x in a for y in b)
  lt =sum((x < y) for x in a for y in b)
 return (gt - lt) / (len(a) *len(b))
 
 
 def compare_group(df, metric, label):
  hpr = df[df.corpus =="hpr"][metric].dropna()
 cmp= df[df.corpus =="comparison"][metric].dropna()
  u, p = stats.mannwhitneyu(hpr, cmp, alternative="two-sided")
  delta = cliffs_delta(hpr.values, cmp.values) iflen(hpr) *len(cmp) <200000elseNone

scipy.stats.mannwhitneyu is a statistical test: it compares two groups of numbers (HPR vs. comparison podcasts, per metric) without assuming they're normally distributed — appropriate for skewed, real-world data like pitch-range measurements. scipy doesn't ship an effect-size measure alongside it, so Cliff's delta — "of all possible (HPR episode, comparison episode) pairs, how much more often is HPR's value the smaller one?" This pairing (a significance test plus an effect-size number) was used throughout because a low p-value alone doesn't say whether the difference is actually big enough to matter.

7. pandas — grouping, stratifying, and cross-tabulating results

print(df.groupby("corpus")["retained_fraction"].describe())
 ...
 print(pd.crosstab(df["corpus"], df["register"]))

Once data/results.jsonl (one JSON object per analyzed episode) is loaded into a pandas.DataFrame, two one-liners answer two different questions that would otherwise need custom code: .groupby(...).describe() gives a distribution summary (count, mean, std, quartiles) of audio-quality coverage per corpus, in one call, and pd.crosstab() cross-tabulates HPR vs. comparison against male-typical/female-typical register — a quick sanity check that the two collections of podcasts had a similar gender-register mix.

8. matplotlib — box plots for the headline comparison

fig, ax = plt.subplots(figsize=(6, 4))
 data = [df[(df.corpus == c) & df[metric].notna()][metric].values for c in ["hpr", "comparison"]]
 ax.boxplot(data, tick_labels=["HPR", "Comparison"], showmeans=True)

A box plot per pitch-range metric, HPR next to comparison podcasts, mean marked as well as the median.

9. spacy word vectors — measuring topic drift, not just word choice

nlp = get_nlp() # spacy.load("en_core_web_md")
 ...
 chunk_docs = [nlp(c) for c in chunks]
 chunk_vecs = []
 for d in chunk_docs:
  v = d.vector
  norm = np.linalg.norm(v)
 if norm >1e-6:
  chunk_vecs.append(v / norm)
 ...
 adj_sims = [float(np.dot(V[i], V[i +1])) for i inrange(n -1)]

Each ~130-word chunk of a transcript is embedded as a single vector using spaCy's medium English model (en_core_web_md, which has built-in word vectors), then L2-normalized so a dot product between two vectors is their cosine similarity. Comparing consecutive chunks' similarity gives a "does the topic barely move between one chunk and the next" score.

10. lexicalrichness — vocabulary diversity that doesn't cheat on transcript length

lr = LexicalRichness(text)
 try:
  mtld = lr.mtld(threshold=0.72)
 exceptException:
  mtld =None

A naive vocabulary-diversity score (unique words ÷ total words) gets worse the longer a text runs, because of how the metric is defined — not because the speaker is actually repeating themselves more.

MTLD (Measure of Textual Lexical Diversity) avoids that by sweeping through the text and counting how many words it takes before diversity drops below a threshold, then averaging — a metric designed specifically to stay comparable between a 10-minute and a 40-minute episode.

11. vaderSentiment — a lexicon-based sentiment score

vader = get_vader() # SentimentIntensityAnalyzer()
 sentiments = [vader.polarity_scores(c)["compound"] for c in chunks]
 ...
 "sentiment_std": float(np.std(sentiments)) iflen(sentiments) >1elseNone,

VADER scores each chunk's sentiment from -1 to +1 using a hand-built lexicon and rule set (not a neural network) — chosen for being fast, deterministic, and cheap enough to run per-chunk across hundreds of episodes. The value then used isn't the average sentiment — it's the standard deviation of sentiment across an episode's chunks: does the emotional

tone ever move, or does it stay flat throughout?

12. zlib (standard library) — compression ratio as a gauge for entropy

def compression_ratio(text):
  raw = text.encode("utf-8")
 iflen(raw) <200:
 returnNone
  compressed = zlib.compress(raw, level=9)
 returnlen(compressed) /len(raw)

No third-party NLP library involved at all — just the standard-library zlib module. A gzip-style compressor finds and exploits repeated patterns; if it can squash a transcript down a lot, that transcript was relatively repetitive and predictable, and if it can barely compress it, the wording was more varied. This is roughly tells you "verbal monotony."

13. re (standard library) — turning "does HPR favour solo hosts" into a checkable heuristic

MULTI_PRESENTER_PATTERNS = [
 r"\bco-?host(?:s|ed|ing)?\b",
 r"\bguest\b",
 r"\binterview(?:s|ed|ing)?\b",
 r"\btalks? with\b",
 r"\bjoins? me\b",
  ...
 ]
 CORRESPONDENT_LINK_RE = re.compile(r'correspondents/(\d+)\.html')
 
 def multi_presenter_score(summary, notes):
  text = (summary or"") +" "+ (notes or"")
  hits = [p for p in MULTI_PRESENTER_PATTERNS if re.search(p, text, re.IGNORECASE)]
  summary_links =set(CORRESPONDENT_LINK_RE.findall(summary or""))
 return {
 "flagged_multi_presenter": bool(hits) orlen(summary_links) >1,
  }

Rather than the acoustic speaker-clustering attempt (item 5, which turned out to be unreliable), a more trustworthy stat came from plain text: HPR's own episode metadata (pulled from a local knowledge-base repo) has host-written summaries, and a short list of regex patterns ("guest", "interview", "joins me", etc.) plus counting linked correspondent profiles was enough to flag likely multi-presenter episodes.

14. multiprocessing.Pool — running the same analysis over hundreds of files without waiting hundreds of times over

work = [(r, local_paths[r["id"]]) for r in batch]
 with mp.Pool(args.workers) as pool, open(args.out, "a") as out_f:
 for result in pool.imap_unordered(_worker, work):
  out_f.write(json.dumps(result) +"\n")
  out_f.flush()

Every batch pipeline (acoustic pitch analysis, speaker-count pass) used the same shape: Pool(workers).imap_unordered(...) to run the per-episode function across multiple CPU cores at once, streaming results out to a JSON-lines file as each one finishes rather than waiting for the whole batch. imap_unordered was chosen over map so a single slow episode couldn't hold up writing out the results that had already finished.

15. subprocess + a resumable batch loop — working within a disk-space limit

def load_done_ids(results_path):
 ifnot os.path.exists(results_path):
 returnset()
  done =set()
 withopen(results_path) as f:
 for line in f:
  ...
  done.add(json.loads(line)["id"])
 return done
 
 ...
 rows = [r for r in rows if r["id"] notin done]
 ...
 for batch_start inrange(0, len(rows), args.batch_size):
  batch = rows[batch_start:batch_start + args.batch_size]
  local_paths = fetch_batch(batch, args.scratch) # scp from the remote host
  ...
 for p in local_paths.values():
  os.remove(p) # delete before next batch

All the podcasts (~220GB) didn't fit on local disk (~54GB free). The fix was a loop: scp (via subprocess) one batch of audio in, analyze it, delete it, move to the next batch. load_done_ids() makes the whole pipeline resumable by reading back what's already in the output file, so a crash or a deliberate stop partway doesn't lose completed stuff.

Links

Ah yes, the riveting world of Hacker Public Radio: a quaint little corner where tech wizards ✨️ divulge their magical secrets ✨️ to each other in the most riveting monotone known to man. Who knew endless "community" podcasts could be so... utterly riveting? ?️?️ Apparently, everyone is just dying to contribute to this cacophony once a year—what a privilege! ?
https://hackerpublicradio.org/
#HackerPublicRadio #techCommunity #podcasting #MonotoneMagic #rivetingContent #HackerNews #ngated


Comments

Subscribe to the comments RSS feed.

Leave Comment

Note to Verbose Commenters
If you can't fit everything you want to say in the comment below then you really should record a response show instead.

Note to Spammers
All comments are moderated. All links are checked by humans. We strip out all html. Feel free to record a show about yourself, or your industry, or any other topic we may find interesting. We also check shows for spam :).

Provide feedback
Your Name/Handle:
Title:
Comment:
Anti Spam Question: What does the letter P in HPR stand for?
Are you a spammer?
Who is the host of this show?
What does HPR mean to you?