Four public feeds, one honest trending table
Shipped
Ghostwriter 0.19.0 rebuilt how the skill researches topics. The audit that forced it was blunt: every scored post sourced from the trending or radar lanes had flopped, six of six, and the “trending” research behind them was a single Hacker News query living in prose, never tuned, never broadened. The release replaces it with trending.py, a deterministic sweep of four public surfaces, plus a computed outcome rollup so the lane rankings come from data instead of vibes. The transferable piece is the sweep itself: how to measure “what’s surging” across several public feeds using only the Python standard library, and how to make the sweep admit when it is broken.
What counts as a signal
The mistake I was making is worth naming before any code: a web search returning articles about a topic is not a trending signal. A signal is a number a surface actually publishes, points, comments, stars, and a recency window. Four public surfaces publish those numbers without any authentication:
| Surface | Endpoint | The signal |
|---|---|---|
| Hacker News | Algolia HN Search API | points, comments, story age |
| Lobsters | /hottest.json |
score, comments, tags |
| Google News | /rss/search with when: |
outlet coverage inside a time window |
| GitHub | repository search API | stars accumulated since creation |
Everything the sweep does is driven by one config file the reader owns. Save this as trending-queries.json and adjust the keywords to your own beat:
{
"interests": [
{
"name": "ai-agents-ops",
"keywords": ["agent", "llm", "sre", "incident", "eval"],
"news_query": "\"AI agents\" (SRE OR DevOps OR incident)"
}
],
"hn": { "min_points": 150, "days": 3 },
"lobsters": { "tags": ["ai", "devops", "practices"] },
"news": { "window": "2d" },
"github": { "topics": ["ai-agents"], "days": 7, "min_stars": 150 }
}
The keywords list is the interest filter: a candidate that matches none of your interests never reaches the table, however loudly it is surging. That rule came straight from the outcome data; high-signal items with no angle I could own were exactly the posts that died.
One fetch, four surfaces
Start trending_sweep.py with a single network wrapper and the two JSON surfaces. Everything here is stdlib; there is nothing to install.
import json, time, urllib.request
USER_AGENT = "trending-sweep/1.0 (research)"
def fetch(url, timeout=15):
req = urllib.request.Request(url, headers={"User-Agent": USER_AGENT})
with urllib.request.urlopen(req, timeout=timeout) as resp:
return resp.read()
def sweep_hn(cfg, get):
hn = cfg.get("hn", {})
cutoff = int(time.time()) - hn.get("days", 3) * 86400
url = ("https://hn.algolia.com/api/v1/search?tags=story&hitsPerPage=30"
f"&numericFilters=points>{hn.get('min_points', 150)},created_at_i>{cutoff}")
return [{
"source": "hn",
"title": hit["title"],
"url": hit.get("url") or f"https://news.ycombinator.com/item?id={hit['objectID']}",
"signal": f"HN {hit['points']} pts / {hit['num_comments']} comments",
"rank": hit["points"],
} for hit in json.loads(get(url)).get("hits", [])]
def sweep_lobsters(cfg, get):
tags = set(cfg.get("lobsters", {}).get("tags", []))
stories = json.loads(get("https://lobste.rs/hottest.json"))
return [{
"source": "lobsters",
"title": s["title"],
"url": s.get("url") or s["comments_url"],
"signal": f"Lobsters {s['score']} pts / {s['comment_count']} comments",
"rank": s["score"] * 15,
} for s in stories if not tags or tags & set(s.get("tags", []))]
Two details carry weight. Every function takes get as a parameter instead of calling fetch directly; that one indirection is what lets tests and frozen baselines replay real captured payloads with no network. And each surface normalizes into the same candidate shape, with a rank on a shared scale. Lobsters scores run far lower than HN points for comparable attention, so the sweep multiplies them; the constant is a judgment call you will tune, not a law.
The HN Search API filters server-side: tags=story excludes comments and polls, and numericFilters accepts points and created_at_i (a Unix timestamp), so the three-day, 150-point floor costs one request. An Ask HN post has no external URL, which is why the item-page fallback exists. Lobsters is simpler: the site’s own Rails routes declare the JSON format for /hottest, and each story carries score, comment_count, and tags you can filter on directly.
News and star velocity
The other two surfaces cover what link aggregators miss: mainstream coverage volume and brand-new projects. Append these to the same file:
import urllib.parse
import xml.etree.ElementTree as ET
def sweep_news(cfg, get):
out = []
window = cfg.get("news", {}).get("window", "2d")
for interest in cfg.get("interests", []):
query = interest.get("news_query")
if not query:
continue
url = ("https://news.google.com/rss/search?q="
+ urllib.parse.quote(f"{query} when:{window}")
+ "&hl=en-US&gl=US&ceid=US:en")
for item in ET.fromstring(get(url)).iter("item"):
out.append({
"source": "news",
"title": item.findtext("title") or "",
"url": item.findtext("link") or "",
"signal": f"News ({window}) · {item.findtext('source') or 'unknown'}",
"rank": 50,
"interest": interest["name"],
})
return out
def sweep_github(cfg, get):
gh = cfg.get("github", {})
since = time.strftime("%Y-%m-%d", time.gmtime(time.time() - gh.get("days", 7) * 86400))
topics = "+".join(f"topic:{t}" for t in gh.get("topics", ["ai-agents"]))
url = ("https://api.github.com/search/repositories?q="
f"created:>{since}+{topics}&sort=stars&order=desc&per_page=10")
return [{
"source": "github",
"title": f"{r['full_name']}: {(r.get('description') or '')[:80]}",
"url": r["html_url"],
"signal": f"GitHub {r['stargazers_count']} stars in <{gh.get('days', 7)}d",
"rank": r["stargazers_count"],
} for r in json.loads(get(url)).get("items", [])
if r["stargazers_count"] >= gh.get("min_stars", 100)]
Google News has no JSON API, but its RSS search endpoint takes the full query syntax including the when: recency operator (when:1h, when:2d, and so on), and xml.etree.ElementTree.fromstring plus iter("item") and findtext are all the parsing it needs. News items get a flat rank of 50 on purpose: an RSS feed has no vote count, recency inside the window is the whole signal, so news should sort below anything with real votes.
For GitHub, created:> plus sort=stars turns the repository search endpoint into a star-velocity query: repos less than a week old, ranked by stars they gathered in that week. A repo at several hundred stars in a few days is surging by any definition.
Rank, dedup, and the empty-sweep rule
The assembly is where two rules live that matter more than any parser. Append this last block:
import sys
SURFACES = {"hn": sweep_hn, "lobsters": sweep_lobsters,
"news": sweep_news, "github": sweep_github}
def match_interest(title, interests):
lowered = title.lower()
for interest in interests:
if any(kw.lower() in lowered for kw in interest["keywords"]):
return interest["name"]
return None
def sweep_all(cfg, get, seen_text=""):
candidates, counts, failures = [], {}, []
for name, sweep in SURFACES.items():
try:
found = sweep(cfg, get)
except Exception as exc:
failures.append(f"{name}: {exc}")
counts[name] = 0
continue
counts[name] = len(found)
candidates.extend(found)
for c in candidates:
c.setdefault("interest", match_interest(c["title"], cfg["interests"]))
fresh = [c for c in candidates
if c["interest"]
and c["url"].lower() not in seen_text
and not (len(c["title"]) > 12 and c["title"].lower() in seen_text)]
fresh.sort(key=lambda c: (-c["rank"], c["title"]))
return fresh, counts, failures
def main():
cfg = json.load(open("trending-queries.json"))
seen = "" # lowercased text of whatever you've already covered
fresh, counts, failures = sweep_all(cfg, fetch, seen)
for failure in failures:
print(f"WARN {failure}", file=sys.stderr)
if not any(counts.values()):
print("ERROR: every surface returned nothing; the sweep is broken, not quiet.",
file=sys.stderr)
return 2
print("raw " + " ".join(f"{k}:{v}" for k, v in counts.items()))
for i, c in enumerate(fresh[:12], 1):
print(f"{i:2}. [{c['source']}] {c['signal']} {c['title']}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Rule one: a surface is allowed to die, and the sweep continues, but the death is printed and counted, never swallowed. Rule two: if every surface returns nothing, that is exit code 2, not an empty table. A sweep with a broken parser and a sweep on a quiet day look identical from the outside; only the raw per-surface counts separate them. In the shipped version those counts also land in a JSON receipt on disk, so later steps can cite exactly what the sweep saw.
The seen_text haystack is the cheapest dedup that works: concatenate everything you have already covered (published posts, recent research notes), lowercase it once, and drop any candidate whose URL or full title appears in it. The length guard on titles exists because short titles substring-match everywhere.
Run it, then starve it
With both files in one directory:
python3 trending_sweep.py
You should see one line of raw counts and a ranked table shaped like this (your items will differ):
raw hn:30 lobsters:11 news:66 github:5
1. [github] GitHub 2116 stars in <7d XiaoDuoYa/codex-with-chatgpt: ChatGPT thinks. Codex works...
2. [hn] HN 390 pts / 117 comments Breaking Claude Code Opus 5 Auto Mode
3. [hn] HN 385 pts / 108 comments I trained a small transformer in 1.5hrs and it beats many LLMs
That sample is from the real run that shipped in this release, and it is the reason the interest filter earns its keep: those 112 raw candidates reduced to 12 rows I could actually use.
Now verify the honest-failure path without touching your network settings, by handing sweep_all a fetcher that always dies:
python3 -c "
import json, trending_sweep as ts
def dead(url, timeout=15): raise OSError('network down')
cfg = json.load(open('trending-queries.json'))
fresh, counts, failures = ts.sweep_all(cfg, dead)
assert counts == {'hn': 0, 'lobsters': 0, 'news': 0, 'github': 0}
assert len(failures) == 4 and fresh == []
print('starved sweep reports itself:', failures[0])
"
Expected output:
starved sweep reports itself: hn: network down
If that assertion holds, your sweep can never silently rot into a green nothing, which is the failure mode that let my old one-curl version go unquestioned for nineteen research runs.
Gotchas
A surface can be unreachable from your network, and only from your network. While building this I planned Reddit as a fifth surface; its JSON listings are widely documented, and both www.reddit.com and old.reddit.com refused the requests from my machine regardless of user agent. The symptom is a JSON decode error, because the body you got back is an HTML block page, not the API. The escape is architectural rather than a workaround: every surface is optional at runtime (the per-surface try/except above), and a surface you cannot verify from where the code actually runs gets cut from the release instead of shipped hopeful. Reddit is deliberately out of this version.
Google News’s query operators are load-bearing and undocumented. The when: operator works today and whole products depend on it, but Google publishes no official reference for it; there is no SLA and no changelog. The symptom when it shifts will not be an error, it will be a feed that quietly ignores your window and hands you stale items. The escape is to treat the feed as a hostile input: parse defensively, keep a captured real payload as a test fixture, and let the raw counts make a sudden shape change visible.
GitHub’s search API rate limit is 10 requests per minute unauthenticated. That sounds like plenty until your sweep grows a query per interest and you run it twice while debugging. The symptom is HTTP 403 partway through a sweep that worked a minute ago. The escape is to design for one search request per run (one combined topic: query, per_page=10) and to remember the failure lands in failures, not on the floor.
Score scales are not comparable across surfaces. Sorting HN points against Lobsters scores raw buries every Lobsters item permanently; the community is smaller, not quieter. The multiplier in sweep_lobsters is the fix, and the honest framing is that any cross-surface constant is editorial. Pick it, write it down, revisit it when the table feels wrong.
Sources
- HN Search API — the Algolia-hosted endpoint, tags and numeric filters
- GitHub REST API: search — repository search,
sort=stars, and the 10 requests/minute unauthenticated limit - Lobsters routes.rb — the JSON format declared for
/hottest - Google News RSS Search Parameters: The Missing Docs — the
rss/searchendpoint andwhen:operator - Google News RSS feeds (FiveFilters) — worked examples of search feeds with recency filters
- xml.etree.ElementTree —
fromstring,iter, andfindtextfor the RSS parsing
Changelog
- ghostwriter 0.19.0 — topic research revamp: measured trending sweep, structural angle gate, computed outcome stats, rebalanced radar (#243) (dbb3fc3)