
How to create skills using screen recordings
I love skills because they turn a useful task prompt into a repeatable way of working. I have written about them in a few editions of this newsletter, and many of you have tried building one since. The sticking point is usually the same: explaining a familiar workflow in enough detail for a model to reproduce it can feel harder than doing the work yourself. The workflows worth the trouble are the complicated ones, twelve steps deep with exceptions hanging off every other step, and writing one down can take longer than doing it.
Anthropic shortened that step last week. Record a Skill turns a screen recording into a draft skill. The explanation happens through demonstration, and instructing an agent moves from writing prompts to showing your screen.
It lives in Cowork, in the Claude desktop app. Start the recording and perform the task while you narrate your decisions. Claude uses the screen activity and your voice to draft the skill. Save it, and a slash command summons it later. (Anthropic says the video and audio are deleted after the skill is built.) A strip of screenshots remains inside the Cowork task.
OpenAI had already introduced a related feature in Codex. Record and Replay captures clicks, typing, and window content. It leaves out narration, which I think carries the more valuable half. Both features suit tasks that are easier to show than describe.
Record a Skill currently works on Pro, Max, and Team accounts, on Mac, inside Cowork. Enterprise accounts cannot use it. A separate screen recording gives those teams another route. I tested that method in May, before the Claude feature existed. Thirteen minutes of QuickTime produced a document with ten sections and screenshots. I used that document to create an SOP, then adapted the instructions into a skill.
The first decision is what belongs on camera. I sort every step of a workflow into five categories: Human in the Loop, Repeatable, Rule or Standard, Context, and Standalone Prompt. Three of those decide the camera question. Repeatable and Rule or Standard steps record beautifully because an explanation for them already exists or could, and the camera is a faster typist than I am. Human in the Loop steps record badly. Their explanation would say "use judgment here," and a recording turns a judgment call into a click with a void behind it.
The camera also generalizes past what it saw. I recorded one instance of an event publishing workflow, a single event with three ticket tiers and two hosts, and the document came back with those exact prices written in as policy, as though a board had ratified them at a summit. Then it advised skipping paid tickets for free events. I never showed it a free event. It extrapolated. Sure, why not.
So record the plainest instance you have, the one you would hand a new hire on day one. Where a workflow forks, record it twice, once down each branch. The free-event advice taught me what the model does with a branch it never saw.
Then talk while you work. Your screen captures what you clicked and typed. Your voice supplies the plan you are following and why one task must happen before another, and that half separates a macro from a skill. Say the exceptions out loud, especially the conditions that never arise during the recording. Mine sounded like "these codes are necessary for our finance team" and "always add tickets are nonrefundable." Talk more than feels natural, even if you sound deranged. The file will not care. The strongest sections of my document covered information that never appeared on screen: the inputs required before starting and the checklist used before publication, and every word of both arrived through the voice track.
Make sure to prepare your desktop before you record, because everything visible on your screen may be captured for the length of the session. I was working inside a live database, so the recording included every record I scrolled past. Close irrelevant tabs and sign out of systems outside the demonstration. Anthropic says deleting the Cowork task will remove the saved screenshots.
Editing turns the recording of one event into a skill that can handle the next one. Test the SKILL.md, the instruction file at the center of the skill, against a case you did not record, and treat every specific value as a suspect, because the model promotes whatever appeared on camera into policy. Replace each promoted choice and restore the general rule behind it. In my event publishing skill, that meant forcing the skill to ask questions at the start of every run: the event, the location, the pricing, instead of assuming my one event was the eternal answer. Mark the steps an agent can run and the decision points that require a person.
There is a real objection here: sitting down to write the document yourself was how you discovered which steps were rules and which were habits. Hand that articulation to a camera and you get a confident draft you have yet to think through, and editing a wrong document can cost as much as composing a right one. A messier workflow could tilt the math, and for me the trade has paid both times.
Audit before you automate, because automation makes whatever it captures permanent. So far mine changed only the skill. The process survived its own audit, and the case for redesigning it now sits in writing, ten sections long.
Record one process that still lives in your head. Sort its steps, clean the screen, record the plainest instance, and say the rules out loud while you click. Let the model draft the skill, then hand it a case the camera never saw. The gaps will show you where the instructions fail and where the process itself needs work.
— Lauren Eve Cantor
thanks for reading!
if someone sent this to you or you haven’t done so yet, please sign up so you never miss an issue.
I’ve also started publishing more frequently on LinkedIn, and you can follow me here
if you’d like to chat further about opportunities or interest in AI, or this newsletter, please feel free to reply.
banner images created with Midjourney.

