← Back to the dashboard

How the Meat Grinder learns

The Grinder decides what to post and when, and it changes those decisions based on what actually happened last time. This page explains how, in plain English — no maths background needed.

The short version

Every post we publish is an experiment. Twenty-four hours later the Grinder looks up how that post did, works out a single score for it, and nudges its own preferences accordingly. Do that a few hundred times and the system's preferences start to reflect what this audience actually responds to, rather than what any of us assumed they would.

Nothing here decides whether something is publishable. A human still approves the content. The learning loop only chooses between takes that have already been approved.

The loop, step by step

  1. A take gets approved. It goes into the Pool — a holding area of approved posts that haven't gone out yet.
  2. The scheduler picks one. Every half hour it looks at what's in the Pool and picks the take it currently believes has the best shot, with some deliberate randomness so it keeps trying things.
  3. It posts, and we wait. The Grinder collects likes, reposts, replies and follower counts at 1, 6, 24 and 72 hours.
  4. The post gets a score. One number summarising how it did — see "How a post is scored" below.
  5. The Grinder updates its beliefs. That score adjusts its opinion of everything the post was made of: the time of day, the graphic template, and the characteristics of the writing itself.
  6. Back to step 2, now slightly better informed.

How a post is scored

Five things get blended into a single score. The first three are about how far a post got; the last two are about what anyone did once it got there.

Engagement rate (30%)
Likes, reposts and replies, divided by how many followers we had when it went out. Dividing by followers makes old and new posts comparable — 200 interactions means something different at 500 followers than at 50,000.
Reach (20%)
Raw interaction counts, not divided by followers. This is the part that notices when a post escaped our usual audience and travelled. Engagement rate alone can't see that, which is why it's here.
Follower change (10%)
How much the account grew in the 24 hours after the post, shared out among the posts that went out in that window — a post that clearly drove the surge gets more of the credit than one that happened to be nearby.
Link clicks (15%)
How many people followed the take's link. Every link we post goes through our own shortener, which counts them. This is the cheapest evidence that somebody didn't just agree, they went and looked.
Conversions (25%)
Signups, donations, RSVPs and volunteer sign-ons on the website, traced back to the specific post that sent them. This is the only part of the score measuring the thing we actually want.

A post carrying no trackable link can't be measured on the last two, so for those the remaining three are simply rescaled to fill the gap. A post is never marked down for a signal it had no way to produce.

Policy Takes are weighted differently again: their whole job is signups, donations and volunteers, so conversions and clicks carry most of their score. And any post a journalist engages with gets a flat bonus on top, because one reporter picking a story up is worth more than almost any post's reach.

Why reach isn't the point

This part changed recently, and it's worth being plain about why. The loop used to be scored on reach alone. Reach is easy to measure, arrives within a day, and is genuinely necessary — a post nobody sees recruits nobody.

But asking people for something costs reach. Every "join us", every petition link, every RSVP is a sentence that isn't a punchline, and the numbers showed it: takes ending on a membership ask consistently reached less far than takes that ended on a joke. Under a reach-only score, that is a finding, and the writer was about to be told about it every week. Follow it and you get an account that travels beautifully and never asks for anything — excellent reach, no members, and nothing on the dashboard showing the mistake.

So the score now measures the whole funnel, and the writer is told to ask for something on a fixed share of takes regardless of what the loop currently believes. That share is not something the optimiser gets a vote on. It costs us some average reach on purpose.

Why scores are mostly small numbers near zero. A post's score is relative to our own recent posts, not an absolute rating. Roughly: zero is a typical post for us, positive did better than usual, negative did worse. A score of −0.3 is not a failure, it's a slightly-below-average Tuesday.

How it decides what to post next

The Grinder keeps a running opinion — with an honest sense of its own uncertainty — about a few different things.

Timing and packaging

For each combination of platform, time slot and graphic template, it tracks how those posts have gone. "Bluesky, 7pm, stat-receipt card", for example.

What the post says

This is the more useful half. Each take gets tagged with about twenty plain characteristics — does it open with a question? Does it name an opponent? Does it contain a dollar figure? Does it end asking people to join? Is it short or long? Does it link out?

The Grinder tracks how posts with each characteristic have performed. Because most takes share characteristics with many other takes, this builds up evidence far faster than tracking individual time slots does — and, importantly, it applies to takes that have never been posted. A brand-new take can be assessed on what it's made of.

Betting on breakouts

Alongside average performance, the Grinder counts how often each option has produced a genuine breakout — a post that reached far beyond our follower base.

This matters more than it sounds. If you only chase good averages, you end up preferring the safe, forgettable post that never flops over the riskier one that lands flat four times and then reaches fifty thousand people. For growth, the second is worth more. So the Grinder deliberately accepts a slightly worse average day in exchange for better odds of a breakout.

Trying things anyway

Roughly one pick in five is made at random, ignoring everything above. Without that, the system would settle early on whatever looked good in its first few weeks and never discover it was wrong. Options it has never tried are also treated optimistically rather than dismissed — an untested idea gets the benefit of the doubt.

Telling the writer what it learned

Measuring performance is only half a loop. The other half is handing those findings back to the step that writes the next take — otherwise the system learns something every week and then writes as though it never had.

So at generation time the model is now given:

Only patterns with enough posts behind them are quoted, never a lucky one-off, and it's phrased as evidence with its sample size attached, explicitly subordinate to the brief — if the story calls for something on the "reached less far" list, the instruction is to write it anyway. A take that misrepresents its source to chase reach is worse than one that underperforms.

…and deliberately ignoring it a quarter of the time

Telling the writer what works has an obvious failure mode: it writes more of what works, the takes start rhyming, and the whole thing slowly converges on one post written forty ways. Worse for our purposes, a system that only polishes its current best will never find the post that would have beaten it — because that post is exactly the one it has stopped trying.

So three things push the other way:

That last one matters more than it looks. The first three all pick from a list of about twenty things we thought to write down and measure. They can find the best combination on that list. They can't find something that isn't on it — and the posts that actually travel are almost never "the right template at the right time". They're a particular sentence that landed. Curveballs are how takes get written that fall outside anything we've catalogued; the measuring machinery then picks up whatever worked, because a genuinely new approach shows up as an odd combination of characteristics that starts winning.

That costs us something — some posts will do worse than if we'd played it safe. That's the trade. An engine that only repeats its greatest hits has stopped looking for the next one.

What it does not do

Reading the dashboard

Analytics → Account Growth
How the accounts are actually doing: followers over time, engagement volume, the furthest-reaching posts, and the funnel underneath them — link clicks, conversions, and press pickups. This is the honest bottom line, independent of anything the model thinks. The pairing worth watching is engagement per post against the click-to-action rate: the first rising while the second falls is the loop buying reach by dropping the ask.
Scheduler → Performance
Whether the generator's own predictions match reality. "Spearman ρ" is just a rank correlation: 1.0 would mean its ranking is perfect, 0 means no relationship, negative means it's backwards. It has been sitting near zero, which tells us that guess isn't currently worth much. The content model figure alongside it estimates how much is predictable from the writing itself — that's the number worth watching.
Scheduler → Arms
The raw state of the loop's beliefs. "Mean" is its current estimate, "pulls" is how many posts that estimate is based on. Anything with only one or two pulls is close to a guess — read those with suspicion.

A caveat worth keeping in mind. The Grinder learns correlations, not causes. If takes about public ownership do well, that might be the topic, or the way we happen to write about that topic, or what else was in the news that fortnight. Treat what it surfaces as a prompt for a human judgement, not a verdict.

The technical decisions behind all of this, including what was deliberately traded away, are recorded in the project's architecture decision records (docs/adr/) — ADR 0006 covers the learning loop's machinery and ADR 0007 covers what it's now pointed at.