privacy

Substack's Data Breach and What Newsletter Platforms Know

In February 2026, a breach exposed 700,000 Substack records. Here's what newsletter platforms collect and what it means for your private writing.

In February, Substack disclosed a data breach that exposed up to 700,000 user records. The records included account details, email addresses, and activity data tied to both writers and readers on the platform. The disclosure came through routine security notifications and received less attention than it deserved — in part because Substack is primarily thought of as a publishing platform, not a data repository.

That framing is worth examining. Substack is also a place where writers store drafts, personal essays, private notes, and the intellectual output of years of work. Readers share their reading patterns, payment details, and email addresses. The platform knows a great deal about both groups, and a breach that exposes 700,000 records is a meaningful reminder that “publishing platform” and “data company” are not mutually exclusive.


What Newsletter Platforms Actually Collect

The category of newsletter and personal publishing platforms — Substack, Ghost, Beehiiv, Buttondown — collects data that most users don’t fully account for when they think of these as simple writing tools.

For writers, the collection typically includes:

  • All content you draft, whether published or not (on hosted platforms)
  • Payment and payout information (bank account details, tax forms)
  • Email list data and subscriber details
  • Engagement metrics: open rates, click rates, unsubscribe events by subscriber
  • Account activity and login history
  • Participation in recommendation networks, which may share subscriber activity across publications

For readers, the collection typically includes:

  • Email address (required to subscribe)
  • Payment information if you subscribe to paid publications
  • Reading activity: which emails you opened, which links you clicked, how long you spent on posts
  • Interaction history with recommendation systems, which feed cross-publication discovery features
  • Device and browser information
  • IP addresses and inferred location

A privacy audit of Substack published in early 2026 gave the platform a score of 40 out of 100 (Grade C). Ghost, which is open-source and can be self-hosted, scored 38 out of 100 (Grade D) in the same audit. Both scores reflect data practices that are standard for the industry but that fall well short of what users typically assume when they think of a personal writing platform.


The Reading Behavior Problem

The most underappreciated element of newsletter platform data collection is reading behavior tracking.

When you open a newsletter email, a tiny invisible image — a tracking pixel — typically registers that event. The platform now knows you opened this specific email, at this time, from this location. When you click a link in that email, the click is routed through the platform’s tracking system before reaching the destination, recording which link you clicked and when.

This tracking is standard across email marketing infrastructure, but it takes on a different character when the emails in question are personal essays, health updates, political commentary, mental health content, or religious writing. The granular record of what you read — and by extension, what you’re interested in, worried about, or thinking through — is a sensitive profile even when the individual data points seem innocuous.

When Substack describes its recommendation network, which surfaces publications to potential subscribers based on reading patterns, the underlying mechanism is this behavioral data. Your reading history isn’t just stored; it actively shapes what the platform shows you and other users. You are both a reader and, in a small way, a signal that influences the platform’s discovery engine.

Writers can opt their publications out of the recommendation network. This prevents their subscriber data from flowing into the cross-publication recommendation system. But reader-level controls over behavioral tracking are more limited.


What the February Breach Exposed

Substack has not published a detailed public breakdown of which specific data elements were included in the 700,000 exposed records. The breach notification language described “user records” and “account information.” Based on what Substack collects (detailed above), exposed records for writers could include subscriber lists, engagement data, and payment information. Exposed records for readers could include email addresses, subscription history, and reading activity.

The most sensitive immediate consequence is email addresses in the hands of unauthorized parties. Email address exposure from a known platform creates direct targeting vectors: phishing emails that appear to be from Substack because the attacker knows you have an account, credential stuffing attempts against your Substack login, and cross-platform targeting using email addresses as identifiers.

If you have a Substack account — as a writer or reader — the practical steps are:

  • Change your Substack password if you haven’t since February
  • Check whether your email address appears in breach notification services
  • Be particularly skeptical of any emails claiming to be from Substack that request credentials or payment details
  • If you use the same password on other services, change it there as well

Ghost: Open Source Doesn’t Mean Private by Default

Ghost is often positioned as the privacy-conscious alternative to Substack — open-source, self-hostable, giving writers full control over their infrastructure. That positioning is partially accurate and partially misleading.

Ghost’s open-source license does mean you can inspect the code. Self-hosting Ghost does mean your content lives on infrastructure you control, not on Substack’s servers. These are real advantages.

The Ghost.io hosted service — where Ghost manages the hosting for you — is a different product. On Ghost.io, your data lives on Ghost’s infrastructure. The privacy score of 38 out of 100 cited above applies to the hosted service. The company still has access to your content, your subscriber data, and your analytics.

Self-hosted Ghost changes this meaningfully, but self-hosting requires technical knowledge: provisioning and maintaining a server, configuring SSL, managing a database, setting up email delivery through a service like Mailgun. The hidden costs — both financial (a VPS plus email delivery typically runs $35–$50 per month) and operational (ongoing maintenance) — are real. For most users, “self-hosted Ghost is more private” is accurate in principle and impractical as a recommendation.


Drafts and Private Writing: More Exposed Than You Think

One dimension of newsletter platform privacy that isn’t discussed often enough is the status of unpublished drafts.

On hosted platforms — Substack, Ghost.io, most newsletter services — content you have written but not published still exists on the platform’s servers. A draft you wrote, reconsidered, and filed away years ago is still accessible to the platform provider, subject to the same legal requests, breach risks, and data handling practices as everything else you’ve stored there.

This matters because people often use personal publishing platforms as a kind of digital journal or extended note-taking system: capturing ideas, drafting through thoughts they never intend to publish, storing longer writing they’re working through over time. The implicit mental model — “this is my private writing, I haven’t published it” — doesn’t match the technical reality. Unpublished means not visible to the public. It doesn’t mean not stored on someone else’s server.

For sensitive drafts, personal essays dealing with health or relationship issues, or writing that touches on topics you wouldn’t want disclosed, the gap between “unpublished” and “private” is worth thinking about explicitly.


What Separation Actually Looks Like

If you want to keep personal writing genuinely separate from platform storage, the options are:

Write locally, publish selectively. Draft in a local application — a text editor, a word processor, a notes app that syncs only to your own devices. Paste content into the publishing platform only when you intend to publish it. This keeps unpublished drafts off the platform’s servers entirely.

Use a local-first notes application for personal writing. Applications like Obsidian store notes on your device and offer sync through infrastructure you control. The trade-off is that they don’t provide email delivery, subscriber management, or the discovery benefits of being on a platform like Substack.

Understand what “private” means on each platform you use. Some platforms offer “members-only” posts visible only to paying subscribers. This limits public visibility but does not limit the platform’s access to the content. “Private from the public” and “private from the platform” are different conditions.

For files, photos, and personal documents you want stored securely — not just inaccessible to other readers but stored encrypted and inaccessible to the platform provider — you need a service specifically built around that model. daftei stores files encrypted in transit using TLS 1.3 and at rest using AES-256, doesn’t analyze file contents, and doesn’t use what you store for advertising or AI training. The purpose is storage that returns your content to you, not platforms that manage your content while also building a business around it.


The Publishing Platform Trade-Off

Newsletter platforms provide real value: distribution infrastructure, subscriber management, payment processing, and discovery through recommendation systems. That value comes with a trade-off that isn’t always clearly stated.

The business model for most hosted publishing platforms — Substack included — depends on their continued operation, growth, and eventual monetization. Your content and subscriber data are part of what they hold and, in various ways, use to sustain that business. The recommendation network is a feature for you and a retention mechanism for them. The reading analytics are presented as tools for writers and also constitute a behavioral profile of readers.

None of this is concealed, exactly — it’s in the terms and privacy policies, in varying degrees of clarity. But the mental model most writers have of a publishing platform (“I write, they distribute, readers read”) undersells how much data the middle layer collects and why.

The February Substack breach is a reminder that this data is a target. 700,000 exposed records represent 700,000 people whose email addresses, account details, and reading histories are now in the hands of whoever found them. For anyone storing sensitive personal writing or maintaining a private subscriber community on a platform like this, it’s a useful moment to reassess what stays there and what doesn’t.

Your memories deserve better than an ad platform.

Try daftei free →
← All posts