How the Galaxy
is made.
Independent research · not affiliated with or endorsed by Substack
2,147 publications form the final influence index. Every one passed the same audience, evidence, influence, relevance and client-suitability rules. Moonlight curation is shown only after admission and cannot change the result.
In briefExecutive summary
We examined all 121,377 public publication URLs listed in Substack’s official sitemap. 12,489 had qualifying first-party evidence of a substantial owned email audience. Of those, 9,863 formed the complete recommendation-analysis frame and 5,261 passed five evidence screens: recent activity, writer control, industry fit, English-language addressability and partnership suitability.
Influence was measured independently across the 9,863-publication graph using owned reach, recommendation authority and cross-community reach. Relevance was then measured from 57 anchors elected by the graph itself: the top 1% by endorsements received within each of the eight industries among the 5,261 five-screen-pass publications. No founder selection is an anchor by virtue of being curated.
The graph produced 3,733 influence-and-relevance candidates. The intersection of those candidates with the 5,261 five-screen publications is the final 2,147-publication index. No target size was imposed.
01Purpose and scope
Culture Galaxy maps independent, writer-controlled publications and the public recommendation relationships between them. It helps brands find editors with meaningful owned audiences, understand which editors carry authority with their peers, and see how attention moves between editorial communities.
“Independent” means the publication is controlled by an identifiable owner-editor and exists primarily for its readers. A writer may run a business or a small editorial team. A corporate marketing list, institutional publication, venture-firm content arm, product newsletter or publication sold through a corporate media channel does not qualify.
The unit of analysis is one canonical public publication URL. Duplicate domains are merged before any count or score is calculated. Private, delisted, newly launched or temporarily unlisted publications cannot be observed through the public sitemap and are outside the stated universe.
The universe is Substack only. Every count, score and ranking here is drawn from Substack's public sitemap. Publications on Beehiiv, Ghost, Kit, Patreon, Mailchimp or self-hosted infrastructure are absent entirely. They are not ranked low or screened out; they are outside the frame. Substack recommendation edges also exist only between Substack publications, so an editor's authority here reflects standing inside one platform's network. This is a map of independent publishing on Substack, and it is not a ranking of independent newsletters generally.
02The complete accounting
| Stage | Count | What the count means |
|---|---|---|
| Public census | 121,377 | Unique publication URLs in the official sitemap. |
| Owned-audience evidence | 12,489 | Direct evidence of 5,000+ email subscribers, or first-party evidence of at least 1,000 paid subscribers when the free count is hidden. |
| Graph-analysis frame | 9,863 | Audience-qualified rows with no recorded hard failure and complete recommendation-source coverage. |
| Five-screen pool | 5,261 | Rows passing activity, writer control, industry, language/market and suitability. |
| Influence-and-relevance candidates | 3,733 | High-influence rows, plus notable rows with high or adjacent graph-derived relevance. |
| Final influence index | 2,147 | The intersection of the five-screen pool and the graph-qualified candidates. |
The last two pools are not independent: graph anchors are elected from the five-screen pool. They are also not subtracted in sequence. The final rule is a logical AND. A publication must pass all five screens and satisfy the influence-and-relevance rule, so 2,147 is the measured overlap, not 5,261 minus 3,733.
03Audience evidence
The primary commercial floor is 5,000 directly evidenced Substack or email subscribers. It is a declared partnership-reach floor, not a claim that 5,000 is a universal scientific threshold. A visible count below 5,000 fails.
When a free count is hidden, an orange or purple Substack Bestseller badge is accepted as a separate owned-audience equivalent. Substack states that these badges represent thousands or tens of thousands of current paid subscribers. The index records that evidence as paid-subscriber evidence; it does not relabel it as a free-subscriber count. Social reach is supporting context only because no defensible universal social-to-email conversion rate exists.
Badge-qualified rows are admitted on equal terms and scored on comparable terms. A bestseller badge supplies a paid-subscriber floor rather than a visible total-subscriber count, so the owned-reach component is imputed rather than measured directly. Each badge-qualified row receives the median measured owned-reach percentile of rows carrying both a measured subscriber count and the same badge tier: 0.890 for “1,000+ paid”, calibrated on 312 rows whose median size is 47,000 subscribers, and 0.985 for “10,000+ paid”, calibrated on 13 rows whose median size is 219,000.
The imputed figure is published in each affected row’s reach field, so every displayed score still reconstructs from its own published components. This approach places badge-qualified publications on the same scoring scale as visible-count publications while preserving the distinction between measured and imputed reach. Two of the 56 named showcase publications rely on this imputation.
Three limits apply. The “10,000+ paid” constant rests on only 13 calibration rows. The calibration rows are by definition those whose subscriber count is recoverable, which may differ systematically from rows whose count is not recoverable in a direction the available data cannot measure. The estimate is a stratum central value, not a measurement: an individual badge-qualified row may sit anywhere in its stratum, whose interquartile range spans 0.780–0.957 for “1,000+ paid” and 0.978–0.993 for “10,000+ paid”. Substituting each stratum’s 25th percentile admits one badge row to the showcase; a single global median of 0.690 admits none. Showcase composition remains sensitive to the imputation constant. Read the authority and bridge components directly where the distinction matters. Measured counts or published intervals would reduce this uncertainty.
Of the 12,489 qualifying rows, 12,331 have direct 5,000+ evidence and 158 qualify through the paid-subscriber badge. A further 22,632 are visibly below 5,000. The unresolved group contains 85,190 hidden-count publications without sufficient evidence, 974 with only a “hundreds paid” badge, one with external reach only and 91 unreachable URLs. These unresolved rows are not described as small; the necessary evidence was simply unavailable.
That missing evidence is not missing at random, and the resulting bias has a direction. Displaying a subscriber count is an editorial choice, and it correlates with running a paid tier, pursuing growth and being comfortable with self-promotion. The 12,489 is therefore not a random sample of publications above 5,000 subscribers: it over-represents commercially oriented editors and systematically omits well-read but publicity-averse ones. A quiet publication with 40,000 subscribers and a hidden count is invisible here, while a smaller publication that displays its number is admitted. Read the index as a map of evidenced and legible owned reach, not of owned reach as such.
A reproducible random sample of 2,000 rows was drawn directly from the unresolved hidden-count population and given a second pass of the same public-page evidence extraction. Of 1,999 reachable pages, two automated text matches were manually rejected as non-audience claims, leaving zero validated evidence recoveries. Because that second pass applies the method these rows already failed, the result is close to circular and is reported as a reproducibility check rather than a population bound. A rule-of-three calculation on it would imply a 95% upper bound near 0.15%, but that figure describes only evidence recoverable by the identical method and is not published as a finding: a genuinely independent recovery estimate would require a different instrument, such as direct publisher contact or archived count snapshots. Their true audience sizes remain unknown. The 12,489 is a strict evidenced lower bound, not a population estimate.
04Five evidence screens
The five screens are separate and collectively cover the non-graph eligibility decision. Their pass counts overlap because one publication can fail more than one screen.
- Recent activity. At least one post in the trailing 60 days. This includes normal monthly publishing while excluding dormant lists. Result across the analysis frame: 9,839 pass, 24 fail.
- Writer control. An identifiable owner-editor, supported by named profile/byline evidence or at least two independent writer signals such as first-person editorial voice, RSS authorship and personal-mode metadata. Corporate, institutional, official, product and branded-author signals fail unless a distinct human editor and writer-controlled relationship are evidenced. Result: 6,230 pass, 3,633 fail.
- Industry fit. One primary industry from the eight-category taxonomy below. Direct publication-owned content evidence controls the assignment. Result: 8,568 pass, 1,295 fail.
- English-language addressability. Predominantly English publication-owned descriptions and recent post titles; nationality is irrelevant. A foreign-language title or name alone never causes exclusion. Missing or contradictory evidence fails closed. Result: 9,530 pass, 333 fail.
- Partnership suitability. Commercially usable editorial inventory for a mainstream client. Result: 8,831 pass, 1,032 fail.
Suitability is conduct-based, not viewpoint-based. Political or conflict reporting, investing analysis, lawful alcohol editorial and medical reporting can qualify. Primary inventory built around overt campaigning, identity antagonism, conspiracy, dangerous treatment claims, explicit adult content, weapons or military-capability promotion, defense procurement advocacy, gambling-like trading signals, exceptional-return promotion or marketing a broker, exchange, fund or financial product does not. Decisions are supported by publication-owned evidence and recorded by reason; keywords flag review but do not make the final decision.
This screen is a brand-partnership filter, and it makes the index narrower than the word “culture” implies. 1,032 publications fail it. Fourteen are transactional crypto, trading and quantitative-finance publications whose primary reader relationship centers on transactions, signals or exceptional-return claims, which conflicts with the suitability rule. Some excluded publications are influential, well-evidenced and widely endorsed; they are excluded because a mainstream advertiser could not buy inventory alongside them, not because they lack standing. The Culture Galaxy therefore maps the commercially partnerable portion of independent publishing. It is not a census of cultural influence, and it should not be cited as one.
04aValidation status and sensitivity
The arithmetic, graph rules, redaction and internal consistency are reproducibly tested. The methodology does not claim a measured real-world accuracy rate for the industry, writer-control or suitability classifications because no independent gold-standard labels or second-rater adjudications exist. Those are declared operational policy screens, not ground truth. Among final industry labels, 60.4% originate in the publication’s Substack directory category, 35.0% in the content-keyword classifier, 3.9% in publisher-description review, 0.6% in prior human audit and 0.1% in named About-page adjudication.
Their load-bearing effect is quantified instead. If each screen were ignored while every other rule remained fixed, the final count would change by at most: activity 0, industry 0, English-language market 15, suitability 390 and writer control 759 publications. These are policy sensitivity ceilings, not estimated mistakes.
Industry labels matter because anchors are elected within industries. Among the 57 graph-elected anchors, 20 carry a keyword-classified label rather than a directory or reviewed label, so classifier error can place an anchor in the wrong industry and shift that industry’s relevance neighborhood.
Remit audit. All 361 Home & Family rows were checked against publisher descriptions and About pages. A second publisher-description review examined 39 flagged rows across all industries and applied all 39 changes. The public audit trail is anonymized: seven moved from Home & Family to Culture & Media, seven from Home & Family to Money & Work, five from Tech & AI to Culture & Media, four from Money & Work to Culture & Media, two each from Home & Family to Health & Self, Style & Beauty to Home & Family and Tech & AI to Money & Work, and one in each of the remaining eleven recorded transitions. City-event coverage was resolved to Travel & Places. A source-name spelling caveat did not affect the writing publication’s clearly evidenced Culture & Media remit. No flag was rejected.
The verified 5,261-row anchor-ranking frame contains industry pools of Money & Work 982, Tech & AI 859, Style & Beauty 374, Food & Drink 511, Travel & Places 116, Home & Family 301, Health & Self 661 and Culture & Media 1,457. Applying the declared allocation elects 57 anchors.
Relevance and membership are computed from the verified taxonomy on the complete 9,863-publication, 35,869-edge analysis graph. Screen-excluded rows remain graph context but cannot become anchors or final members. Hops, relevance bands, candidate status and final membership are propagated across the full graph, producing 3,733 research candidates and a 2,147-publication final index. Industry enters screening, anchor election and relevance, but not authority, bridge or owned reach.
Location-population audit. Locale is populated only when the publication’s own description or About page explicitly states a current base location. Regions, inferred locations, topic locations, past locations and ambiguous multi-city bases are left blank. Borough and common abbreviation forms are normalized to city names. Locale is populated for 96 of 2,147 rows across 37 distinct city values; the complete city counts are published in the downloadable method metadata.
Placeholder-description review. The two flagged “My personal Substack” rows were retained because each independently passes all five frozen evidence screens. Cadence remains calculated from the crawl ending August 1, 2026, while the edition banner is dated August 11, 2026; apparent cadence-versus-latest contradictions are snapshot timing, not row-level date errors.
One structural conflict is disclosed rather than mitigated. The curator is also the sole classifier, and Moonlight & Company has a commercial interest in the network the index describes. Every discretionary judgment (industry, writer control and suitability) is single-rater, made by an interested party, with no second adjudicator. The build controls for this mechanically by excluding the curation annotation from every formula and by electing anchors without it, and the checks above quantify what remains. It is not eliminated, and readers should weigh discretionary labels accordingly.
05Influence
Influence is computed for every row in the 9,863-publication analysis frame before relevance or curation is considered. It uses three independent metric families:
- Recommendation authority. 65% PageRank-style authority plus 35% source-normalized endorsement authority. Source normalization divides each recommender’s unit of authority across its outgoing list, so unusually long lists cannot dominate. These correlated measures are combined and counted once.
- Cross-community reach. A participation measure that rises when a publication connects several recommendation communities rather than remaining inside one, adjusted for the volume of unique relationships.
- Owned reach. Verified audience evidence on a logarithmic scale, so size matters without allowing the largest lists to overwhelm the graph evidence.
A publication is high influence when at least one family is at or above the 95th percentile, or at least two are at or above the 80th. It is notable when at least one is at or above the 80th, or at least two are at or above the 65th. All others are developing. These classes are mutually exclusive.
The displayed 0–100 score is 45% recommendation authority, 25% cross-community reach and 30% owned reach. It is a transparent ranking aid. Admission uses the percentile rules above, not an unpublished score cutoff. The full download publishes all three analysis-frame percentile components as authority, bridge and reach, so the rounded display score can be independently reconstructed.
Substack Recommendations is a cross-promotion feature, not a disinterested citation system, and the graph inherits that. Editors enable recommendations partly to acquire subscribers, and reciprocal arrangements are a common and openly discussed growth tactic on the platform. In the 9,863-row frame, 16,250 of 35,869 directed edges (roughly 45%) belong to mutual pairs, a rate consistent with substantial reciprocity. Recommendation authority should therefore be read as measuring peer esteem and participation in a promotional network together; the two cannot be separated with public data alone. An editor who declines to use the feature is under-measured here regardless of the regard peers hold them in.
06Graph-derived relevance
Relevance is graph-derived: anchors are elected from public recommendation patterns within each industry, not from Moonlight’s curated selections.
Within each of the eight industries, the method elects the top 1% of five-screen-pass publications by unique endorsements received. Network authority and canonical URL break ties deterministically. The 5,261-row ranking frame produces 57 graph-elected anchors: Money & Work 10, Tech & AI 9, Style & Beauty 4, Food & Drink 6, Travel & Places 2, Home & Family 4, Health & Self 7 and Culture & Media 15. Section 04a records the reviewed classifications and completed propagation. Curation flags are withheld from election and scoring; thirteen anchors overlap the curated set only when the annotation is joined back afterward.
A publication is high relevance within one recommendation step of a graph anchor. It is adjacent within two steps or when its detected recommendation community contains at least two anchors. Everything else is contextual. These bands are mutually exclusive and exhaustive.
High-influence publications qualify globally. A notable publication must also be high or adjacent. This retains important outliers while preventing a merely mid-level graph score from entering solely because it is large or topically similar.
Relevance is computed on the complete 9,863-publication analysis graph. Within the final index, 10,948 relationships have both endpoints in the index; that count is published as totals.edges. The restricted preview ships only the 110 directed relationships running between the 56 named publications, so no withheld publication appears as an edge endpoint. Either way the displayed constellation is an induced viewing subgraph and cannot by itself reproduce analysis-frame hops, communities or bridge scores.
07Communities and categories
Recommendation communities are detected from the weighted graph without demographic labels or curator seeds. Nine random-seed and resolution runs are compared. The reference partition is the highest-modularity result among the three runs at the declared 1.0 resolution; the other six resolution runs are sensitivity checks. The selected partition has modularity 0.585; median adjusted-Rand agreement is 0.543 and the minimum is 0.473. In the final index, 309 rows have community stability below 0.5. Ten admissions use the two-anchors-in-community clause; communities otherwise position the map rather than override the influence rule.
The content taxonomy is mutually exclusive and exhaustive for the final index: Money & Work, Tech & AI, Style & Beauty, Food & Drink, Travel & Places, Home & Family, Health & Self, and Culture & Media. Each row receives one primary industry from its own content evidence. Secondary topics such as careers, entrepreneurship, faith or politics remain searchable subject matter inside the primary buying lens; suitability is evaluated separately.
The final industry counts are Culture & Media 673, Money & Work 364, Tech & AI 354, Health & Self 231, Food & Drink 231, Style & Beauty 167, Home & Family 95 and Travel & Places 32. The connection bands are also mutually exclusive: 1,109 core, 941 adjacent and 97 contextual. Industry answers “what does it cover?” Community answers “who recommends whom?” Connection answers “how close is it to graph-elected market hubs?” Influence answers “how much authority and reach does it carry?” None is used as a substitute for another.
08Curation and presentation
Seventy final-index publications carry the ✦ sparkle because they were already known to Moonlight or sourced through women in Tennessee’s network. The sparkle is an annotation added after every eligibility decision. It cannot supply audience evidence, create an anchor, improve a score, change a category or admit a publication. Publications carrying the sparkle are expected to be concentrated among influential editors because they were originally chosen for visible peer esteem; the build verifies qualification with the annotation field excluded from every formula.
White stars are the top 1% of final influence scores, brass stars are the next 19%, and the remaining 80% form the surrounding field. These three tiers are assigned once, by rank on the rounded displayed score, and published per row as tier; canonical URL breaks exact score ties. The brass boundary falls inside a three-way tie at 85.1. Canonical URL places Book Enthusiast at rank 430 in brass, while Hungry Dogs with James Patterson and The GameDiscoverCo newsletter begin the field. The map reads the published tier field rather than re-deriving tiers from preview score bands. Territory planets label industries only. Constellation lines show the final-index subset of public recommendation relationships in either direction when a publication is selected.
The restricted preview names exactly seven publications per industry: the highest final influence scores, with canonical URL as the tie-break. This creates 56 showcase identities. The same rule is applied after full qualification in every industry; curation, logo availability and founder preference are not selection inputs. Because the rule ranks on the displayed score, it inherits that score’s imputation limits: two of the 56 are badge-qualified rows whose owned reach is an imputed stratum median rather than a measured count, and section 03 records how sensitive that membership is to the imputation constant. Its downloadable data publishes only relationships among those 56 named publications and replaces exact locked-publication metrics with broad bands; headline totals still describe the full audited index.
Eleven of those 56 showcase identities carry the ✦ sparkle, compared with 70 of 2,147 in the full index. This concentration is an observed result of the top-seven influence ranking, not a quota or override, and is disclosed at the preview’s conversion boundary.
The directory defaults to Recommending publications, the raw count of unique analysis-frame publications recommending that editor. “Overall influence” sorts the composite score; “Cross-community reach” sorts the bridge measure. These are ranking views of the same qualified index, not extra admission tests.
09Freshness
The Signal checks a selected set of high-influence publication feeds and refreshes hourly when a live reading is available. Otherwise it displays a clearly dated snapshot. Opening a publication card checks that publication’s latest post at that moment when possible. Missing, empty, invalid or pre-2020 refresh dates are omitted rather than rendered as epoch dates. Cadence is the median number of days between observed RSS posts. The Cadence filter’s “Weekly” band and the headline weekly-or-more count both mean a median of seven days or fewer; on that definition 1,300 of the 1,746 publications with a measured cadence qualify. The publication card uses a deliberately looser wording, “weekly-ish”, for medians up to ten days. Withheld preview rows publish no cadence and are filed under “Withheld in preview” rather than counted as unmeasured.
Audience evidence, recommendations, ownership, category, suitability, scores and membership change only through a complete audited census refresh. This prevents a transient response from silently changing the headline denominator. A publication that crosses the audience floor or gains network authority can enter when all stages are rerun together.
10Identity marks and data handling
Publication identity uses 504 stored publication marks, 963 distinct favicons, 416 publication-controlled page marks and 266 Moonlight crescents where no usable mark is evidenced. An unavailable image falls back to the crescent; generic globes and generic platform marks are not intentionally presented as publication identities.
The research uses publicly visible metadata from RSS feeds, About pages, recommendation pages, the public directory and the official sitemap. It does not access subscriber lists, private accounts or paywalled article text. Stored fields are factual metadata: publication name and address, feed title and description, explicitly stated base city, post titles, links and dates, cadence, audience evidence and recommendation relationships. Article bodies are not retained or redistributed.
Publication names and marks are used for identification and link directly to their owners. Collection is limited to publicly served pages and feeds and does not use credentialed or paywalled access. Where a platform or publisher term requires a specific arrangement for recurring collection, that arrangement is to be confirmed before the next census rather than assumed.
11Limits
- The universe is Substack only. Publications on other platforms are absent from every count and ranking, and authority measures standing within one platform's recommendation network.
- The public sitemap cannot reveal private, delisted, newly launched or temporarily omitted publications.
- The 85,190 unresolved hidden-count rows are unknown, not declared ineligible, and they are not missing at random: displaying a subscriber count correlates with commercial orientation, so the index over-represents growth-minded editors and under-represents publicity-averse ones of equal reach. The recovery audit is a reproducibility check, not a population bound.
- Badge-qualified publications receive an imputed owned-reach value (the median of measured rows sharing their badge tier) rather than a measured count. Two of the 56 named publications rely on that imputation, the “10,000+ paid” constant is calibrated on only 13 rows, and substituting each stratum’s 25th percentile changes which badge rows are named. See section 03.
- Substack Recommendations is a growth feature with a reciprocity incentive; 45% of directed edges are mutual. Authority measures esteem and promotional participation together, and editors who do not use the feature are under-measured.
- Public recommendations are evidence of editorial endorsement, not proof of readership, agreement, demographics or purchase intent.
- The partnership-suitability screen excludes 1,032 publications on brand-safety and reader-relationship grounds, including 14 transactional crypto, trading and quantitative-finance publications. The index maps commercially partnerable independent publishing, not cultural influence at large.
- Older publications have had longer to accumulate recommendations. The 60-day activity test removes dormant incumbents but does not make the graph age-neutral.
- Editors tend to recommend similar editors. The derived, cross-industry anchor rule reduces founder bias, but thirteen of the 57 anchors overlap the Moonlight network, so curation could still influence admission through anchor proximity. No observational network can eliminate homophily.
- 35% of graph-elected anchors carry keyword-classified industry labels with no measured accuracy rate. The 361-row Home & Family remit audit and 39-row publisher-description review verify labels against first-party evidence, but industry-level membership remains correctable evidence rather than immutable ground truth.
- Discretionary labels are single-rater, produced by a party with a commercial interest in the result, with no second adjudicator.
- The analysis graph contains 35,869 directed endorsements. The final downloadable subgraph contains 10,948 directed edges representing 8,303 unordered publication pairs; 138 final publications have no visible final-to-final edge. Their influence may derive from recommendations involving analysis-frame publications outside the final index.
- Gender, age, income and “multi-hyphenate” identity are not inferred. Audience fit for a specific client remains commissioned research.
