Skip to content

About Subdex

Public data is difficult to explore

Reddit’s history is scattered across independent archives, each with its own coverage, freshness and query language. Answering a straightforward question — what did this community discuss in 2021 — usually means learning an API, writing a script, and discovering the gaps by accident. The data is public. Working with it is not straightforward.

Research, not identity hunting

There is a category of tool that takes a username and promises to reveal a person. Subdex is deliberately not that. It analyzes public content and stops there: no location inference, no identity resolution, no cross-site matching, no personality scoring, no judgement about whether two accounts are one person.

Those absences are the product. A tool that can answer a research question without also being an instrument for finding someone is more useful to the people doing legitimate work, and less useful to everyone else.

Multiple sources

Records come from independent archives that disagree with each other. Rather than hiding that behind a single confident number, Subdex labels which source returned what and preserves conflicts where they occur. Disagreement between archives is information about coverage, not noise to be cleaned up.

A local workspace

There is no account and no server-side database. Analysis runs in your browser; saved research lives in your browser. Nothing is uploaded, which also means nothing is recoverable if you clear your browser data — a real tradeoff, stated plainly rather than glossed over.

Honest about coverage

Archive search is useful precisely because it is imperfect, so the interface should show you where the gaps are. Every figure says whether it came from an archive or was calculated locally, partial results are marked as partial, and score statistics exclude records captured before their scores had settled.

The alternative — a clean number with no provenance — is more comfortable and considerably more misleading.