How I Make Viral Faceless AI Videos Step by Step

You don’t need to show your face, hire a camera crew, or know how to edit video to build a YouTube channel that gets real views. Alston runs multiple faceless channels in niches like remote jobs, painting tutorials, and language learning, and in this walkthrough he lays out the exact step-by-step process he follows, including the specific tools, the exact prompts, and the money-versus-time trade-offs you’ll face at every turn.

This is not a theoretical framework. It’s the actual workflow behind channels that have crossed monetization thresholds, picked up tens of thousands of views on single videos, and kept running without Alston ever appearing on camera. The process costs money to run at speed, but the logic behind it is simple: find what YouTube has already proven it wants to show, and remake it in a way that’s genuinely yours.

What You’ll Walk Out With

  • The exact method for finding videos that are already going viral in any niche
  • How to use Poppy AI to reverse-engineer why a video performs, then generate a unique script from that analysis
  • Why voice cloning in HeyGen beats using template AI voices, and how to set it up
  • The Adobe Premiere Pro + Storyblocks B-roll workflow that keeps editing time to 15-20 minutes
  • A ChatGPT prompt sequence for SEO titles, descriptions, tags, and thumbnail ideas, all from a single transcript
  • How to use Leonardo AI and Photoshop to finish a click-optimized thumbnail in under 10 minutes
  • The honest cost breakdown: what’s paid, what’s skippable, and when free tools make sense
  • Ready to find the right niche for your faceless channel? Start at finder.platformproof.com

Step 1: Build a Viral Video Research List

The entire process starts with research, and the research is more systematic than most people expect. Alston begins by finding 20 to 30 channels that are already covering his target niche. He doesn’t filter by size. He looks at channels with a million subscribers and channels with 3,000 subscribers sitting side by side in the same spreadsheet.

From those channels he’s looking for two things. First, videos that have gone viral in the last month, which he defines as videos with more views than the channel has subscribers. A channel sitting at 5,000 subscribers that just posted a video with 20,000 views is a strong signal. That video is getting pushed by the algorithm to people who don’t follow that channel, which means the concept is proven. Second, he’s looking at all-time best performers across the channel’s history. He uses Shane Hummus as a benchmark example, a creator in the remote jobs space who has multiple videos sitting at 5 million views despite only having around 1 million subscribers. Those are the targets.

All of it goes into a Google Sheet. Two columns: the channel, and the video URL. That list becomes the raw material for everything that follows.

Step 2: Run the Video Through Poppy AI

Poppy AI is the centerpiece tool in this workflow. It’s a paid service, there is no free tier, and Alston describes it as a step above ChatGPT because it can actually transcribe a YouTube video from a URL, something ChatGPT can’t do as of this recording.

The prompt he uses is direct: paste the video URL and ask Poppy AI to analyze why the video goes viral. Poppy AI transcribes the content, then breaks down the structural elements that drove retention, including hook construction, pacing, the way information is layered, and what keeps viewers from clicking away. The output isn’t a summary of what the video is about. It’s an explanation of how the video works as a retention engine.

From there, Alston asks Poppy AI to turn that analysis into a framework he can hand to a VA. This is important: the goal at this stage is not a finished script. It’s a blueprint that captures the structural logic of what made the original video perform.

Step 3: Generate a Modified Script

With the framework in hand, Alston goes back to Poppy AI and asks it to write a 20-minute script, but with a twist. He modifies the original topic so the resulting video is genuinely different from the source material.

His example: if the original viral video is titled “8 Certifications to Find a Remote Job,” he doesn’t make that same video. He makes “8 Certifications to Find a Remote Job for Accountants,” or for nurses, or for logistics workers. He adds a qualifier that narrows the audience, which makes the content feel more specific and targeted to a real subgroup while still riding the structural proof that the broader concept works.

The first draft Poppy AI produces often runs long, sometimes 30 minutes of spoken content. Alston asks it to rewrite until the script lands between 15 and 20 minutes. That’s the target length for the kind of mid-format video that earns ad revenue and ranks well in search without losing viewers who expected a short watch.

This is where he addresses the ethics question directly. YouTube doesn’t reward novelty. It rewards proven ideas. What Alston is doing is not copying someone’s video. He’s using structural insight from a successful video to build something new. The information is different, the angle is different, the script is original. Taking what’s proven and making it your own is how most of YouTube works, even at the top of the platform.

Step 4: Generate Audio and Video With HeyGen

Once the script is locked, Alston copies the full text and pastes it into HeyGen. You can also use ElevenLabs for audio-only generation, but his preference is HeyGen because it simultaneously produces both an audio track and a video of the AI avatar speaking. He’ll delete the video layer later and keep only the audio, but having the video lets him review timing before committing to the edit.

The specific thing he emphasizes here is voice cloning. He’s subscribed to a HeyGen tier that has cloned his own voice, so the output sounds like him, not like a generic AI voice. His rule is firm: do not use template voices. Audiences have learned to recognize stock AI narration, and it kills watch time. If you’re going to run a faceless channel, the audio needs to sound like a real person with a specific voice, not a text-to-speech reader from 2019.

HeyGen takes time to render because it’s processing both audio and video at once. While it renders, you move on to the next step instead of waiting around.

Step 5: Edit in Adobe Premiere Pro With Storyblocks B-Roll

When the HeyGen file is ready, Alston imports both the audio and video into Adobe Premiere Pro. He immediately deletes the video track and keeps only the audio. What he’s building is an audio-driven video where the visual layer is entirely B-roll.

For B-roll, he uses the Storyblocks plugin that integrates directly into Premiere Pro. He searches for footage relevant to whatever topic the script covers, for example a video about remote accounting jobs would pull in footage of accountants, spreadsheets, home offices. He downloads clips directly into the Premiere Pro project through the plugin, then drags and drops them onto the timeline.

His editing rule for B-roll is a clip change every 3 to 5 seconds. This keeps the visual pace moving without feeling frantic. He renders everything in 4K. The entire edit including the B-roll search and the drag-and-drop work takes roughly 15 to 20 minutes when the source material is already organized. These days he has a video editor handling this step for him, but the process is the same.

Adobe Premiere Pro has an auto-transcription tool built in, and Alston uses it to generate captions automatically. His caption settings: all caps, Impact font, white text on a black background. This is a deliberate choice. That caption style is native to social media short-form content, and audiences trained on Reels and Shorts are comfortable with it. It increases retention because viewers can follow along even without sound. After captions are applied, he hits render and moves immediately to the next step while the file processes.

Not sure which niche is right for your faceless channel?

Answer a few questions and find out at finder.platformproof.com.

Step 6: Use ChatGPT to Build the Entire Upload Package

While Premiere Pro is rendering, Alston opens ChatGPT and downloads the transcript that Premiere Pro generated. He pastes that transcript into ChatGPT and runs a single prompt that asks for five SEO-optimized title options, a description, timestamps, and tags, all at once.

ChatGPT processes it in about 5 to 10 seconds and returns a full upload package. He copies the output, pastes it into YouTube Studio, and picks the title he wants. The description and timestamps go in as-is or with minor tweaks. This step, which most creators spend 30 to 60 minutes on, takes under two minutes.

From the same ChatGPT session, he runs a second prompt: using this title and description, give me 10 thumbnail ideas optimized for click-through rate. He then asks ChatGPT to pick the one it believes gives the highest probability of a strong CTR and to write a detailed image generation prompt for that thumbnail that he can paste into Leonardo AI.

Step 7: Create the Thumbnail With Leonardo AI and Photoshop

Alston takes the thumbnail prompt from ChatGPT and pastes it into Leonardo AI. He generates two versions of the image so he has options. Leonardo AI handles the text and compositional elements based on the prompt. He downloads both versions to his computer.

The final step is Photoshop. He opens the downloaded image and makes three adjustments only: increase brightness, increase contrast, increase vibrance. That’s the entire Photoshop workflow. These three adjustments make the thumbnail pop against a gray YouTube feed background, which is the only environment where the thumbnail actually needs to compete. He saves it and uploads it to YouTube.

That’s the complete process from niche research through live upload.

The Honest Cost of Running This System

Alston doesn’t pretend this workflow is free. Here’s what it actually costs to run the full version:

  • Poppy AI: Paid subscription, no free trial available
  • HeyGen: Paid, with a tier that includes voice cloning (the feature you actually need)
  • Adobe Premiere Pro: Paid, part of the Adobe Creative Cloud subscription
  • Adobe Photoshop: Paid, bundled with most Creative Cloud plans
  • Storyblocks: Paid subscription for the Premiere Pro plugin and stock footage library
  • Leonardo AI: Paid, no functional free tier for thumbnail volume

Alston’s framing on this is direct. You have either time or money. If you have more time than money right now, there are free alternatives at almost every step: free AI voice tools instead of HeyGen, free editing software instead of Premiere, free stock footage instead of Storyblocks. The paid tools compress the time each step takes. The free tools trade that time back in exchange for zero upfront cost.

The channel that proved the model for him in the remote jobs niche has run profitably across multiple videos. His painting channel crossed the YouTube Partner Program threshold and has earned $55 from ad revenue, which he acknowledges is lunch money but also proof that the monetization mechanism works. The Spanish words channel crossed 1,000 subscribers using a different version of the same basic approach: find a narrow topic, produce consistent content at volume, let the algorithm figure out who to show it to.

The Niche Expansion Play: Spanish Words and Language Channels

One of the more creative examples Alston shares is his language learning channel. He built it by hiring a native Spanish speaker through Fiverr to record individual words and short phrases. The video format was simple: Alston narrated the intro and context in English, the Fiverr narrator said the Spanish phrase correctly, and the phrase repeated three times. He doesn’t speak Spanish. He failed Spanish in high school and actually took Chinese. That didn’t stop him from building a channel about it.

The channel has crossed 1,000 subscribers with around 1,700 hours of watch time, meaning it still needs more hours to hit the full monetization threshold. He’s exploring whether to pivot it toward travel content, where the Spanish connection still makes sense, or whether to go back to its roots and just produce longer Spanish language lesson videos, around 5 minutes instead of the original 1-minute format.

He’s also tested translating existing videos into Hindi using HeyGen’s auto-translation feature. That channel has 6 subscribers at time of recording, but the point he’s making is about what’s possible with the tools. HeyGen can take a finished English audio track and produce a translated version in another language with consistent voice characteristics. For a language learning channel, that’s a content production multiplier, not just a one-off novelty.

The monetization angles for a language channel include a paid course teaching the language, a digital download like a vocabulary guide for travelers, or affiliate partnerships with language learning platforms. The channel builds the audience; the audience is the asset.

Step-by-Step Summary: The Full Faceless AI Video Process

  1. Research 20-30 channels in your target niche. Build a Google Sheet of videos with high views-to-subscriber ratios (last 30 days) and all-time best performers on those channels.
  2. Paste one of those video URLs into Poppy AI. Ask it to analyze why the video goes viral and produce a framework you can hand to a VA or use yourself as a creative brief.
  3. In the same Poppy AI session, use that framework to generate a 15-20 minute script on a slightly modified version of the original topic. Add a qualifier, change the audience, narrow the angle.
  4. Paste the full script into HeyGen (or ElevenLabs for audio-only). Use your cloned voice, not a template voice. Let it render.
  5. Import the HeyGen file into Adobe Premiere Pro. Delete the video track, keep the audio. Add B-roll from Storyblocks at a clip change rate of every 3-5 seconds. Export in 4K.
  6. Apply auto-captions in Premiere Pro: all caps, Impact font, white text on black background. Render.
  7. While rendering, download the Premiere Pro transcript and paste it into ChatGPT. Get 5 title options, an SEO description, timestamps, and tags in one prompt.
  8. In the same ChatGPT session, get 10 thumbnail ideas, ask it to pick the best one for CTR, and generate a detailed Leonardo AI image prompt.
  9. Run the prompt in Leonardo AI. Generate 2 versions. Download both.
  10. Open the thumbnail in Photoshop. Raise brightness, contrast, and vibrance. Save. Upload everything to YouTube.

Find Your X

The system works across niches because the logic is the same regardless of topic. Find what’s already working, understand why it works, make a version that’s genuinely yours. The question most people get stuck on isn’t the tools or the steps. It’s figuring out which niche to start in and which angle within that niche gives them the best shot at gaining traction before running out of motivation.

That’s what the Platform Proof Finder is built for. Answer a few questions about your background, your available time, and your goals, and it points you toward the specific direction that fits. Start at finder.platformproof.com.

Frequently Asked Questions

Do I need to show my face to make money on YouTube?

No. Alston runs multiple profitable faceless channels where the content is entirely AI-generated audio over B-roll footage. The key is that the voice sounds like a real specific person, not a generic AI narrator, which is why he uses voice cloning instead of template voices in HeyGen.

Is Poppy AI worth paying for?

Alston uses it as his primary AI tool for video research and scriptwriting because it can transcribe a YouTube URL directly, which ChatGPT cannot do. Whether it’s worth it depends on how many videos you’re producing per week. If you’re making one video a month, the cost-to-output ratio is harder to justify. If you’re doing multiple videos per week across multiple channels, it pays for itself in time saved.

What’s the difference between HeyGen and ElevenLabs for this workflow?

HeyGen produces both audio and a synchronized AI avatar video, which Alston finds useful for reviewing timing before editing. ElevenLabs produces audio only. Both support voice cloning. Alston uses HeyGen because the full audio-plus-video output gives him more to work with in the edit, but he deletes the video portion anyway and keeps only the audio track.

Isn’t this method just copying other people’s videos?

No, and Alston addresses this directly. He’s taking structural insight from a successful video, not the content itself. The script is new, the angle is modified, the information serves a different subaudience. The approach is similar to how a book publisher studies what made one business book a bestseller and then commissions a different author to write a new book on a related angle. The structure is learned from; the content is original.

How long does the full process take per video?

The B-roll editing step alone takes about 15-20 minutes once you have your audio file. The ChatGPT and Leonardo AI steps take under 10 minutes combined. The biggest time sink is the Poppy AI research and scripting phase, which can take anywhere from 30 to 90 minutes depending on how many iterations the script needs to hit the right length. Total time with the paid tool stack is roughly 2-3 hours per video, with a lot of that being render wait time during which you can work on something else.

What niches work best for faceless AI channels?

Alston has tested remote jobs, painting tutorials, Spanish language learning, Hindi language content, and is exploring others. Common traits among the niches that worked: people are actively searching for specific answers, the content can be broken into list formats or step-by-step guides that work well with B-roll, and there are already channels proving that the audience exists. Evergreen search topics outperform news-driven topics for this model because they keep getting discovered long after the upload date.

What’s the minimum budget to start this system?

Alston doesn’t name exact dollar amounts in this video, but the tool stack includes Poppy AI, HeyGen (with voice cloning tier), Adobe Premiere Pro, Adobe Photoshop, Storyblocks, and Leonardo AI, all of which are paid subscriptions. If budget is a constraint, the most important paid tool to prioritize is the one that clones your voice, because that single element has the biggest impact on watch time. Everything else has a free or cheaper alternative, even if it takes more time.

Can this system work if I don’t have expertise in the niche?

Yes, and Alston’s Spanish channel is the clearest proof of that. He openly says he failed Spanish in high school. He built the channel by hiring a Fiverr narrator to say the phrases correctly. The channel’s value to viewers came from the format and the accessibility of the information, not from Alston’s personal expertise in Spanish. The same principle applies when Poppy AI is generating scripts based on proven successful content: you’re curating and structuring knowledge, not inventing it from personal experience.

Read Next

If this process resonated with you, the natural next question is what to do when a faceless channel isn’t growing as fast as expected. The follow-up post breaks down the specific mistakes that stall faceless channels and how to fix them without starting over.

How to Fix Your Faceless YouTube Channel (Stop Losing Money!)

Sources

  • Alston Godbolt, “How I Make Viral Faceless AI Videos Step by Step,” YouTube, 2024, https://youtu.be/yME9u4w2rpA
  • Poppy AI (poppyai.com) – AI tool for YouTube video transcription and viral analysis
  • HeyGen (heygen.com) – AI video and voice cloning platform
  • ElevenLabs (elevenlabs.io) – AI voice cloning and text-to-speech
  • Adobe Premiere Pro – professional video editing software
  • Storyblocks (storyblocks.com) – stock footage subscription with Premiere Pro plugin
  • Leonardo AI (leonardo.ai) – AI image generation for thumbnails
  • Adobe Photoshop – used for final thumbnail brightness/contrast/vibrance adjustments

Related Reading


Helping 1 million working adults make their first $3,000 online with the skills they already have. Alston Godbolt, Platform Proof.