@plugdash/fromsubstack

fromsubstack

1,050 words·5min read

What it does

Registers Substack as a source in EmDash's admin importer. Upload the export ZIP you get from Substack's Settings > Exports, and fromsubstack reads posts.csv plus each post's HTML file, converts the body to Portable Text, and reports images so EmDash's importer can bring them in too.

It does not write any content itself. EmDash's importer calls into this plugin to preview and stream posts, then does the actual creating. There is no CLI step - the whole thing runs off the uploaded ZIP, in memory.

Install

npm install @plugdash/fromsubstack

Register

Open astro.config.mjs at the root of your Astro project. Import the plugin at the top, then add it to the pluginsarray of the emdash() integration - the one that sits inside Astro's top-level integrations array:

astro.config.mjsjs
import { defineConfig } from "astro/config"
import emdash from "emdash/astro"
import { fromsubstackPlugin } from "@plugdash/fromsubstack"

export default defineConfig({
  integrations: [
    emdash({
      plugins: [
        fromsubstackPlugin({
          targetCollection: "posts",
          status: "draft",
          importImages: true,
          preserveSlugs: true,
        }),
      ],
    }),
  ],
})

Once registered, "Substack" appears as a source in EmDash's admin importer. status only decides what happens to posts Substack itself had marked published - a post that was still a draft on Substack always imports as a draft, no matter what you set here. That default is deliberate: review what came across before any of it goes live.

Import your Substack

Download your export

In Substack, go to Settings > Exports and request an export. Substack emails you a link to a ZIP containingposts.csv (one row per post), an HTML file per post, and an attachments folder of images that posts reference by relative path.

Upload the ZIP

Open EmDash's admin importer and pick "Substack" as the source. Upload the ZIP exactly as Substack gave it to you - don't unzip it first, or fromsubstack won't find posts.csv at the root.

Review the analysis

Before anything is imported, EmDash previews the export: how many posts it found, how many distinct images they reference, and whether your target collection's schema can take the import. A collection needs a required title (string) and body(portableText) field, plus optional subtitle (string) andpublished_at (datetime) fields. If the schema doesn't fit, the preview names which field is missing or the wrong type.

Run the import

Confirm the import and EmDash streams the posts in, pulling image bytes for anything it decides to keep. Posts that were still drafts on Substack come in as drafts regardless of the statusoption; posts Substack had published come in as draft or published depending on how you set it.

What maps to what

For each posts.csv row that has a matching HTML file:

SubstackEmDash
titletitle
subtitleexcerpt, and meta.substackSubtitle
post_date / date / published_atdate
the post's HTML filecontent, converted to Portable Text
first image found in the bodyfeaturedImage
slug from the post's URL (or the title, if preserveSlugs: false)slug
post_id / idsourceId, and meta.substackId
urlmeta.substackUrl
audience columnmeta.substackAudience, and meta.substackPaid (true unless it's everyone)
is_published: falsestatus forced to draft, regardless of config
is_published: truestatus follows the status option (draft or publish)
type podcast or threadskipped - no article body to convert
row with no matching HTML fileskipped, with a warning
a slug that collides with one already importedskipped, with a warning

The body conversion (headings, bold/italic, links, images, blockquotes, fenced code with language detection, and lists) is handled by@plugdash/html-to-portable-text. Tags it doesn't recognize are unwrapped - their text stays, the tag doesn't.

A worked example

Say posts.csv has this row:

csv
post_id,post_date,is_published,type,audience,title,subtitle,url
101,2024-01-05T10:00:00Z,true,newsletter,everyone,"Hello, World",The very first one,https://acme.substack.com/p/hello-world-jan

paired with posts/101.hello-world-jan.html:

html
<h2>A heading</h2>
<p>Plain text with <strong>bold</strong> and <em>italic</em> and
<a href="https://example.com">a link</a>.</p>
<img src="https://substackcdn.com/image/one.png" alt="One">

With the default options, that becomes an EmDash item shaped like this:

yaml
sourceId: "101"
postType: "post"
slug: "hello-world-jan"
title: "Hello, World"
excerpt: "The very first one"
date: 2024-01-05T10:00:00.000Z
status: "draft"          // default status option is "draft"
featuredImage: "https://substackcdn.com/image/one.png"
content: [ /* h2 heading, a paragraph with bold/italic/link marks, an image block */ ]
meta: {
  substackId: "101",
  substackUrl: "https://acme.substack.com/p/hello-world-jan",
  substackAudience: "everyone",
  substackPaid: false,
  substackSubtitle: "The very first one"
}

A row with audience: only_paid or foundingcomes across the same way, except meta.substackPaid istrue. EmDash has no native idea of a Substack paywall, so that flag is left for you to act on - gate the post yourself, or pair the import with something like @plugdash/paygate if you need the paywall to carry over.

Known limitations

  • Doesn't create content on its own - it's a source for EmDash's importer, which does the actual writing.
  • Doesn't guarantee pixel-perfect HTML - unknown tags are unwrapped and their text kept, not preserved as markup.
  • Doesn't resolve Substack's paywall into an EmDash paywall - paid and free posts both import, just flagged differently in metadata.
  • Doesn't deduplicate against posts already in EmDash from an earlier, different import - only duplicate slugs within the same export are caught.
  • Doesn't download images during the preview step - it reports their URLs, and only fetches bytes for the ones the importer decides to keep.
  • Doesn't import podcast episodes or other non-article post types - Substack's export lists them, but there's no body to convert.
  • Paid status reflects the export's audience column, not a live check against Substack's subscriber list.

Props

proptypedefaultdescription
targetCollectionstring"posts"Collection posts import into
status"draft" | "published""draft"Status for posts Substack had marked published (a post still a draft on Substack always imports as a draft)
importImagesbooleantrueReport body images so EmDash imports them
preserveSlugsbooleantrueKeep the slug from the Substack post URL instead of re-slugifying the title
onWarn(message: string) => voidconsole.warnWhere per-post warnings (skips, empty bodies) go

For agents

After installing @plugdash/fromsubstack and registering it in astro.config.mjs:

  1. Confirm the plugin shows up as "Substack" in the admin importer's source list.
  2. Have the user download their export from Substack (Settings > Exports) and upload the ZIP there.
  3. Review the analysis: post count, attachment count, and whether the target collection's schema can take the import (schemaStatus.canImport). If it's false,schemaStatus.reason names the missing or mismatched field.
  4. Run the import. Metadata written per post: meta.substackId,meta.substackUrl, meta.substackAudience,meta.substackPaid, and meta.substackSubtitle.
  5. If a post is missing after import, check the warnings the host surfaced - a duplicate slug, a CSV row with no matching HTML file, or a non-post type are all skipped rather than failing the whole import.
built with plugdash.dev