News sentiment score of headline + summary, averaged per day (−1 to +1)
Sentiment mix by category
Share of negative / neutral / positive headlines
NegativeNeutralPositive
Topic hierarchy
Category › theme › topic. Articles counted once per node; click a label to filter, the caret to expand
Sentiment by topic
Mean news sentiment per topic (topics with at least 10 articles)
Most-covered stories
Articles clustered by event (same news across outlets, plus different stories about the same event). Click a story to see its coverage
Story
Outlets
Articles
Sentiment
Topic path
Span
Duplicate detection
How much of the selection repeats an earlier article, at each level
People
Most-mentioned persons (spaCy NER). Click to filter
Organisations
Most-mentioned organisations. Click to filter
Places
Most-mentioned countries, cities and regions. Click to filter
Headlines
Hover a sentiment score to see which phrases produced it.
Published
Headline
Source
Category
Sentiment
Topics
Entities
Coverage
Source health
Source
Category
Region
Status
Entries
New
Latency ms
Last runs
Note
Built by python -m newscollector report. Sentiment: finance/politics phrase lexicon with negation, intensifier and contrast handling (newscollector/sentiment.py) blended with VADER, on headline + summary; entities: spaCy en_core_web_sm; topics: keyword rules and hierarchy in newscollector/nlp.py; duplicates: exact / near / same-event clustering in newscollector/dedupe.py, scored against a hand-labelled gold set on every run. Scores are indicative, not editorial judgements.
Episodes and their sections
Each bar is one episode split into its topical sections (width = share of the episode's words), coloured by the category of the section's main topic. Hover a section for its topics, click an episode to list its sections
Topics over time
Share of each month's discussion per topic (section words × topic score). Click a row to filter
Topic hierarchy
Category › theme › topic. Sections counted once per node; click a label to filter, the caret to expand
People
Sections mentioning each person (hosts excluded). Click to filter
Organisations
Labs, companies and institutions by sections mentioning them
Products & models
Models, products, tools and protocols by sections mentioning them
Sections
Episode date
Section
Words
Topics
Entities
Tone
How episodes were sectioned
Where each episode's section boundaries came from
Built by python -m newscollector podcast + report. Episodes and transcripts from the publication's Substack API (newscollector/podcast.py); sections from the author's chapters, transcript headings or automatic TextTiling (newscollector/segment.py); entities: spaCy en_core_web_sm plus an AI gazetteer and model-name patterns; topics: AI keyword rules and hierarchy; tone: the news sentiment scorer averaged over speaker turns (newscollector/podcast_nlp.py). Scores are indicative.