Peter Yang walks through the ChatGPT Skills chain that runs his podcast production — 15 manual steps a week, three skills, one orchestrator. (19 minutes, his own channel; prompt pack at behindthecraft.com.)

The three-step system

  • Map every manual step. He wrote out all 15 things he used to do per episode — research the guest, build the interview guide, send edit instructions, transcript, upload, title/thumbnail, newsletter, clips, social posts. Just making the list exposed three distinct phases hiding inside it
  • Build one skill per phasepodcast prep, podcast edit, podcast production. Production is the orchestrator: it walks him through thumbnail, title, show notes, newsletter, social post and clips one by one, calling the individual skills
  • Connect the skills end to end. His framing: if you find yourself repeating the same template or process, build a skill for it

The first move after the list is to paste the list into ChatGPT and ask “what work from here can you take off my plate?” He runs that in a custom mode so his existing skills don’t shape the answer — he wants to see what the model claims before it sees his setup.

Prep — research that can kill the episode

Real example: an episode with Ethan from OpenAI on ChatGPT for personal finance.

  • One prompt (“I’m interviewing Ethan about personal finance, please research and prep an interview guide”) triggers the podcast prep skill
  • Output: research links, an end-goal for the interview, example practical use cases, and a YouTube scan of what performs in the niche
  • That scan is the point — if there are no interesting videos in the space, maybe the interview isn’t worth recording at all. It also proposes title/thumbnail packaging up front
  • He then iterates: look up the ChatGPT finance docs and Ethan’s tweet history, restructure the guide around what ChatGPT Finance is → connecting accounts → use cases. The revised guide ships to Ethan as a Google Doc

Edit — the handoff to humans

  • The podcast edit skill opens a Linear ticket for the video editors, then builds the editor package
  • Intro reel: 30–40 seconds of the most engaging moments, drafted as two options from the written transcript, which he then asks to be combined into the version he actually wants
  • Clip candidates pulled from the transcript with social copy attached (e.g. “audit nine days of spend”), then iterated with start/stop feedback
  • Moments to cut, found by scanning the whole episode — logistic issues, technical difficulties — with timestamps. This is the piece he says would cost him hours manually: listen to the whole podcast again and pick out the moments
  • He uses the Riverside MCP inside ChatGPT rather than the app’s UI to drive this

Production — packaging, where the time actually goes

  • Newsletter postpodcast production calls podcast post for a first draft. The first draft wasn’t good; he told it which six prompts from Ethan to incorporate, pasted them in verbatim, then made manual edits himself
  • Thumbnail + title — a skill that re-reads the transcript, finds the main topics, researches outperforming videos on his channel and similar ones, and returns five packaging combinations. He counters with his own pairings and they converge. Even with AI it’s still 30–40 minutes per episode because he’s picky — and once the thumbnail template is stable, a Figma MCP (or whatever tool) can edit it directly
  • Show notes — YouTube and Spotify descriptions with real timestamps from the uploaded video, sponsor links, newsletter link, and the guest’s links, which it found on its own. Refinement: name the actual use cases in the timestamps. Then it drives the browser to paste the copy into the YouTube description, toggle monetization, and set the playlist
  • Social post — two teasers and two podcast posts from the podcast social skill (built from examples of his past posts). He picks the intro quote, then asks it to find a screen cap from the video where they’re both smiling — it finds it, attaches it, and the post is scheduled on Typefully
  • Clips — asks for top clip suggestions, picks the one with the best banter, makes it through the Riverside MCP, cuts pauses and filler, reviews, and schedules the clip with its copy

The technique worth stealing: diff your own edits back into the skill

The most reusable part of the video isn’t a skill file, it’s the loop he uses to improve them.

  • Before editing the AI’s newsletter draft by hand, he asks it to snapshot the current post
  • He makes his manual edits
  • He asks it to compare the two and list what changed, then summarize the before/after into a numbered list of changes to fold into the skill
  • He reads and reviews the list, gives feedback, gets a revised list, and says “make all the updates to the skill”

That’s how the same mistakes stop recurring. Same pattern for show notes and clips: name what broke, have the model write the rule, then commit the rule into the skill.

What stays human

  • He’s in every step — reviewing output, editing by hand, directing the clips
  • Skill files look intimidating, but he didn’t write most of it: he gave instructions and a workflow, and the file carries when to use it and how
  • He sticks to a mid-tier GPT model rather than the most expensive option
  • More connected MCPs and plugins means more of the workflow can be automated
  • His own admission at the end: even with the system, producing the podcast still eats his time, and he probably should hire a human to drive the AI instead of doing all of it himself

If you don’t have a podcast

  • Pick one or two manual workflows that eat your week
  • List every step, without skipping the annoying ones — more context and instructions means more of it gets taken
  • Look for the phases that emerge from the list, then have AI build a skill for each
  • Connect them and automate as much as you can

I’m still really involved in each step to review AI’s output, to make manual edits if necessary, because I don’t want it to just pump out slop. I want it to be high taste and high craft.