fromsubstack
What it does
Registers Substack as a source in EmDash's admin importer. Upload the export ZIP you get from Substack's Settings > Exports, and fromsubstack reads posts.csv plus each post's HTML file, converts the body to Portable Text, and reports images so EmDash's importer can bring them in too.
It does not write any content itself. EmDash's importer calls into this plugin to preview and stream posts, then does the actual creating. There is no CLI step - the whole thing runs off the uploaded ZIP, in memory.
Install
npm install @plugdash/fromsubstackRegister
Open astro.config.mjs at the root of your Astro project. Import the plugin at the top, then add it to the pluginsarray of the emdash() integration - the one that sits inside Astro's top-level integrations array:
import { defineConfig } from "astro/config"
import emdash from "emdash/astro"
import { fromsubstackPlugin } from "@plugdash/fromsubstack"
export default defineConfig({
integrations: [
emdash({
plugins: [
fromsubstackPlugin({
targetCollection: "posts",
status: "draft",
importImages: true,
preserveSlugs: true,
}),
],
}),
],
})Once registered, "Substack" appears as a source in EmDash's admin importer. status only decides what happens to posts Substack itself had marked published - a post that was still a draft on Substack always imports as a draft, no matter what you set here. That default is deliberate: review what came across before any of it goes live.
Import your Substack
Download your export
In Substack, go to Settings > Exports and request an export. Substack emails you a link to a ZIP containingposts.csv (one row per post), an HTML file per post, and an attachments folder of images that posts reference by relative path.
Upload the ZIP
Open EmDash's admin importer and pick "Substack" as the source. Upload the ZIP exactly as Substack gave it to you - don't unzip it first, or fromsubstack won't find posts.csv at the root.
Review the analysis
Before anything is imported, EmDash previews the export: how many posts it found, how many distinct images they reference, and whether your target collection's schema can take the import. A collection needs a required title (string) and body(portableText) field, plus optional subtitle (string) andpublished_at (datetime) fields. If the schema doesn't fit, the preview names which field is missing or the wrong type.
Run the import
Confirm the import and EmDash streams the posts in, pulling image bytes for anything it decides to keep. Posts that were still drafts on Substack come in as drafts regardless of the statusoption; posts Substack had published come in as draft or published depending on how you set it.
What maps to what
For each posts.csv row that has a matching HTML file:
| Substack | EmDash |
|---|---|
title | title |
subtitle | excerpt, and meta.substackSubtitle |
post_date / date / published_at | date |
| the post's HTML file | content, converted to Portable Text |
| first image found in the body | featuredImage |
slug from the post's URL (or the title, if preserveSlugs: false) | slug |
post_id / id | sourceId, and meta.substackId |
url | meta.substackUrl |
audience column | meta.substackAudience, and meta.substackPaid (true unless it's everyone) |
is_published: false | status forced to draft, regardless of config |
is_published: true | status follows the status option (draft or publish) |
type podcast or thread | skipped - no article body to convert |
| row with no matching HTML file | skipped, with a warning |
| a slug that collides with one already imported | skipped, with a warning |
The body conversion (headings, bold/italic, links, images, blockquotes, fenced code with language detection, and lists) is handled by@plugdash/html-to-portable-text. Tags it doesn't recognize are unwrapped - their text stays, the tag doesn't.
A worked example
Say posts.csv has this row:
post_id,post_date,is_published,type,audience,title,subtitle,url
101,2024-01-05T10:00:00Z,true,newsletter,everyone,"Hello, World",The very first one,https://acme.substack.com/p/hello-world-janpaired with posts/101.hello-world-jan.html:
<h2>A heading</h2>
<p>Plain text with <strong>bold</strong> and <em>italic</em> and
<a href="https://example.com">a link</a>.</p>
<img src="https://substackcdn.com/image/one.png" alt="One">With the default options, that becomes an EmDash item shaped like this:
sourceId: "101"
postType: "post"
slug: "hello-world-jan"
title: "Hello, World"
excerpt: "The very first one"
date: 2024-01-05T10:00:00.000Z
status: "draft" // default status option is "draft"
featuredImage: "https://substackcdn.com/image/one.png"
content: [ /* h2 heading, a paragraph with bold/italic/link marks, an image block */ ]
meta: {
substackId: "101",
substackUrl: "https://acme.substack.com/p/hello-world-jan",
substackAudience: "everyone",
substackPaid: false,
substackSubtitle: "The very first one"
}A row with audience: only_paid or foundingcomes across the same way, except meta.substackPaid istrue. EmDash has no native idea of a Substack paywall, so that flag is left for you to act on - gate the post yourself, or pair the import with something like @plugdash/paygate if you need the paywall to carry over.
Known limitations
- Doesn't create content on its own - it's a source for EmDash's importer, which does the actual writing.
- Doesn't guarantee pixel-perfect HTML - unknown tags are unwrapped and their text kept, not preserved as markup.
- Doesn't resolve Substack's paywall into an EmDash paywall - paid and free posts both import, just flagged differently in metadata.
- Doesn't deduplicate against posts already in EmDash from an earlier, different import - only duplicate slugs within the same export are caught.
- Doesn't download images during the preview step - it reports their URLs, and only fetches bytes for the ones the importer decides to keep.
- Doesn't import podcast episodes or other non-article post types - Substack's export lists them, but there's no body to convert.
- Paid status reflects the export's
audiencecolumn, not a live check against Substack's subscriber list.
Props
| prop | type | default | description |
|---|---|---|---|
targetCollection | string | "posts" | Collection posts import into |
status | "draft" | "published" | "draft" | Status for posts Substack had marked published (a post still a draft on Substack always imports as a draft) |
importImages | boolean | true | Report body images so EmDash imports them |
preserveSlugs | boolean | true | Keep the slug from the Substack post URL instead of re-slugifying the title |
onWarn | (message: string) => void | console.warn | Where per-post warnings (skips, empty bodies) go |
For agents
After installing @plugdash/fromsubstack and registering it in astro.config.mjs:
- Confirm the plugin shows up as "Substack" in the admin importer's source list.
- Have the user download their export from Substack (Settings > Exports) and upload the ZIP there.
- Review the analysis: post count, attachment count, and whether the target collection's schema can take the import (
schemaStatus.canImport). If it'sfalse,schemaStatus.reasonnames the missing or mismatched field. - Run the import. Metadata written per post:
meta.substackId,meta.substackUrl,meta.substackAudience,meta.substackPaid, andmeta.substackSubtitle. - If a post is missing after import, check the warnings the host surfaced - a duplicate slug, a CSV row with no matching HTML file, or a non-post type are all skipped rather than failing the whole import.