Blog
Notes on working with Reddit archives: what they support, where they mislead, and how to tell the difference.
Digital Research
A Better Way to Search Years of Public Reddit Comments
Comment search is the slowest thing an archive does. Working with that constraint beats fighting it.
2026-08-22 · 3 min read
Reddit Research
A Practical Introduction to Researching Public Reddit Discussions
A workflow for getting from a vague question to a defensible finding, and the places it usually goes wrong.
2026-08-22 · 3 min read
Digital Research
A Researcher's Guide to Reddit Timestamps
Timestamps look like the simplest field in the data. They are where reproducibility quietly breaks.
2026-08-22 · 3 min read
Data Analysis
Activity Heatmaps: Useful Visualization or Easy to Overinterpret?
The calendar grid is persuasive out of proportion to what it shows. That is exactly the problem.
2026-08-22 · 3 min read
Digital Research
Building Reproducible Research From Public Reddit Archives
Archives change underneath you. Reproducibility means recording enough that someone can tell what you had.
2026-08-22 · 3 min read
Public Archives
Deleted, Removed, and Archived: Three Different Things
Three states that look identical in a result list and mean completely different things. Confusing them produces confident, wrong conclusions.
2026-08-22 · 3 min read
Online Communities
Posts vs Comments: Two Very Different Signals
Treating posts and comments as one category of activity throws away the most useful distinction in the data.
2026-08-22 · 2 min read
Research Ethics
Public Online Research Without Crossing the Line Into Doxxing
The line is not public versus private. It is whether your work assembles a person or describes a pattern.
2026-08-22 · 3 min read
Reddit Research
Reddit Flair as Research Data: Useful, but Easy to Misread
Flair is one of the few structured fields Reddit offers. It is also community-specific, unverified, and frequently a joke.
2026-08-22 · 3 min read
Research Ethics
The Ethics of Studying Deleted Public Content
Deletion is the clearest signal a person can send about their own content. Archives preserve it anyway.
2026-08-22 · 3 min read
Data Analysis
The Problem With Treating Archived Scores as Final Numbers
A score in an archive is a measurement taken at one moment. Some of those measurements are placeholders.
2026-08-22 · 3 min read
Reddit Research
Understanding Reddit Comment Trees
A comment quoted alone can mean the opposite of what it meant in its thread. Structure is not decoration.
2026-08-22 · 2 min read
Data Analysis
Using Keyword Frequency Without Pretending It Reveals Personality
Word counts describe text. The leap from text to person is where word-frequency analysis usually goes wrong.
2026-08-22 · 3 min read
Online Communities
What Makes a Subreddit Activity Timeline Useful?
A rising line can mean growth, a moderation change, an archive gap closing, or Reddit itself growing. Usually you cannot tell which.
2026-08-22 · 3 min read
Public Archives
What Reddit Archives Can — and Cannot — Tell You
Reddit archives answer some questions well and others badly. Knowing which is which is most of the skill in using them.
2026-08-22 · 3 min read
Online Communities
What Subreddit Participation Can Tell You — and What It Cannot
Where someone posted is a fact. What it means about them is an inference, and usually a bad one.
2026-08-22 · 2 min read
Public Archives
Why Archive Sources Sometimes Disagree
Two archives of the same site return different answers. The disagreement is data, not an error to resolve.
2026-08-22 · 3 min read
Digital Research
Why Local-First Research Tools Are Useful for Public Data
If the data is public, why does it matter where the analysis runs? Because the query is not the same thing as the data.
2026-08-22 · 3 min read
Public Archives
Why Reddit Archive Search Results Are Never Perfect
Incompleteness is not a bug in archive search. It is a property of the problem, and the interface should show you where it bites.
2026-08-22 · 3 min read
Data Analysis
Working With Large Reddit Datasets in the Browser
Fifty thousand records is fine in a browser. The things that make it slow are rarely the things people optimise.
2026-08-22 · 3 min read