ContentBeginner to intermediate

AI Content Creator: 15 Practical Modules

A practical curriculum from briefs and prompts to images, video, voiceover, avatars, animation, music, asset rights, QA, and a portfolio.

10 min readReviewed Sep 9, 2026Free public access
Table of contents

This course takes you from a simple brief to a reviewable content package: text, images, video, voiceover, avatars, animation, virtual try-on, and music. The goal is not to use the most tools. It is to build a process whose output, cost, assets, permissions, and revisions can be traced.

Examples use a fictional Kedai Pagi business and a fictional adult character named Nara. Replace them with your own topic. Never upload another person's face, voice, work, or data without appropriate authority and consent.

How to Use the Course

Create this working structure:

text
01_Brief/
02_References_and_Permissions/
03_Prompts/
04_Images/
05_Audio/
06_Video/
07_Edit/
08_Final/

For every attempt, record the objective, input, tool/model, meaningful settings, cost or credits, result, problem, and next decision. Features, menus, quotas, prices, and licences change; check official terms before buying or publishing.

Complete Modules 1-5 first. Then choose the story/animation path (8-10) or personal-presenter path (6 and 13). Modules 7, 11, 14, and 15 are specialist tracks. Apply the cost discipline from Module 12 from your first generation.

Module 1 - AI Fundamentals

Outcome: distinguish ideas, prompts, models, references, generation, editing, and human evaluation.

Confident AI output can still be wrong. Content is complete only after a person checks facts, context, quality, and usage rights.

Practice: a testable brief

text
Task: create a 30-second vertical video concept.
Audience: young professionals struggling to start the day.
Message: one small step is better than waiting for perfection.
Fixed facts: Kedai Pagi and Nara are fictional; no discounts or testimonials.
Format: time, visual, narration, and on-screen text table.
Constraints: one character, one location, three shots.
Done when: one message is clear and every shot is producible.

Evidence: one-page brief and a list of facts to verify.

Module 2 - Fast, Clear Prompts

Use six parts: Task, Context, Input, Constraints, Format, Success criteria. Longer is not automatically better; every sentence should support a decision or evaluation.

Create versions A, B, and C. Add one constraint at a time, compare outcomes, and do not change every variable together.

text
Create three content concepts for Kedai Pagi for young professionals. Use only facts in the attached brief. Do not invent prices, discounts, real locations, or testimonials. Return a table with hook, message, three scenes, CTA, and facts still needing confirmation. Use one character and one location; each concept must fit three shots.

Evidence: three prompt versions, outputs, and the reason for your final choice.

Module 3 - AI Images and Character Consistency

Build a character bible: adult/fictional status, face, hair, accessories, clothing, proportions, rendering style, and three details that must not change. Approve one master reference and reuse it.

text
Create a reference sheet for Nara, a fictional adult. Show front, three-quarter, side, and one natural expression. Keep short wavy hair, navy apron, cream shirt, plain round pin, and a soft 3D illustration style. Neutral background, no text, no change in age or outfit.

Change one variable per test. Simplify difficult hand poses and add exact typography during editing.

Evidence: character bible, master reference, four scenes, and face/outfit/product QA.

Module 4 - AI Video

Use text-to-video for exploration, image-to-video for animating an approved frame, and an editor for assembling shots, audio, text, and transitions. Image-to-video prompts should emphasize subject, camera, and environmental movement.

TimeVisualPurpose
0-5sShop door and morning lightEstablish place
5-12sNara places a cupMain action
12-20sCup, notebook, deskEnding and title space
text
Animate this reference for 5 seconds. Nara slowly moves one cup, stops, then looks at the table. Locked camera, small natural movement, stable morning light. Do not change the face, outfit, fingers, cup shape, or written marks.

Evidence: storyboard, three clips, one 20-second edit, and a failure log.

Module 5 - Text to Speech and Voiceover

Write for listening: one message, short sentences, punctuation for pauses, pronunciation notes, and one primary tone. Time the script by reading it aloud.

text
Turn this brief into a 35-second Indonesian voiceover. Use warm, informative spoken language. Preserve every fact, invent nothing, mark pauses with line breaks, and list words whose pronunciation must be tested. Output the script, then the pronunciation list.

Generate or record the voice, clean noise, and keep background music below speech on phone speakers.

Evidence: final script, clean voice, mixed version, and pronunciation notes.

Module 6 - Your Own Face in Images and Video

A reference-based portrait, an animated photo, and a digital twin are different processes. Use your own identity or a consenting adult. Check whether clothing or background falsely implies a profession, endorsement, or testimonial.

Create one well-lit source image, record permission and purpose, generate one portrait, compare identity details, and animate a short test. Follow the provider's consent and verification process for a digital twin.

Evidence: source, permission, portrait, short video, and identity/motion/transparency QA.

Module 7 - Image and Video Face Swap

Face swap combines source identity, target media, and a mask/blend. One preview frame cannot validate a video. Review every frame around head turns, occlusion, lighting changes, and multiple faces.

Use only owned or licensed source and target assets. Test an image first, then a very short clip. Do not use public figures, children, victims, or deceptive contexts.

Evidence: source/target rights, two simulations, and frame-by-frame QA. Technical smoothness alone does not make publishing safe.

Module 8 - Story Content

A story needs causal change, not just attractive images. Use opening, desire, obstacle, action, change, and ending.

text
Create a 60-second story synopsis about Nara feeling overwhelmed and choosing to tidy one table. The character and business are fictional. Do not frame it as a testimonial or therapy advice. Return a logline, six causal beats, narration, shot list, emotion, props, and continuity notes.

Lock script and audio before long renders. Make the riskiest shot early so credits are not spent on unusable scenes.

Evidence: synopsis, six beats, storyboard, script, and understandable rough cut.

Module 9 - Animation Content

Choose one style: flat 2D, paper cutout, soft 3D, or illustrative stop motion. Movement can be small when it supports the message. Lock the character, style, ratio, background, and final audio. Start with a fixed camera and one action; reserve clean space for editor-added captions.

Evidence: style frame, character reference, short animation, and cross-clip consistency notes.

Module 10 - Natural Lip Sync

Cartoon animation and realistic avatars need different inputs and carry different risks. Use clean speech without music. If audio changes after processing, realign the timeline.

Check audio start, names/numbers/consonants, mouth movement during music-only sections, and framing. Fix one broken word with a targeted audio or clip revision instead of rerendering everything without diagnosis.

Evidence: test clip, problem timestamps, targeted revision, and final output.

Module 11 - AI Virtual Try-On

Virtual try-on is a concept visualization, not a promise of size, fit, material, or exact product detail. Use a licensed model image and a complete product image.

Compare colour/pattern, collar/buttons, sleeves/length, logo/text, and body proportions. Mark each result concept, needs revision, or ready for review. Use real photography when material product details cannot be preserved.

Evidence: model/product inputs, three results, QA table, and visualization disclaimer.

Module 12 - Tokens, Credits, and Local Alternatives

Input/output tokens, media credits, rate limits, context windows, and licences are different constraints. Free, trial, and local do not automatically mean unlimited, zero-cost, commercially safe, or compatible with your hardware.

AttemptRevision objectiveCost/creditsPassed?Decision
S01Stabilize camera......Keep/change/stop

After three attempts with the same symptom, stop. Write the diagnosis, change one variable, or use manual editing/capture when it is more reliable.

Evidence: budget, attempt log, and cloud/local/manual decision rules.

Module 13 - Voice Cloning and Face Avatars

Voice cloning models a voice identity; a face avatar models visual identity. Both need consent, access control, purpose, usage period, and a clear withdrawal plan. Provider verification still applies.

Use only your own identity for this exercise. Record a clean sample, test a short phrase, and evaluate similarity, pronunciation, stability, model privacy, and publishing context.

Evidence: sample, consent record, access setting, audio test, optional avatar, and appropriate AI disclosure.

Module 14 - Songs and Music Videos

Use original lyrics and themes. Define mood, genre, instruments, vocal type, structure, and approximate tempo without asking to imitate a specific performer or song.

text
Create a direction for an Indonesian acoustic-pop song about starting the day with one small step. Warm and optimistic; acoustic guitar, light piano, simple bass, soft percussion; natural adult vocal; memorable chorus; around 96 BPM. Do not imitate any artist or song. I will provide original lyrics separately.

Put the final audio on the timeline first, mark real verse/chorus times, and storyboard visual progression. A singing character is optional for a beginner music video.

Evidence: lyrics, music prompt, final audio, licence notes, storyboard, and a 30-60 second clip.

Module 15 - Covers and AI Vocal Conversion

Separate composition/lyrics, master recording, instrumental, performer, and voice-model rights. Voice conversion changes timbre; it does not fix pitch or grant song rights. Attribution, disclaimers, stem separation, and AI labels do not create permission.

Use your own composition, lyrics, vocal recording, a licensed instrumental, and your own voice model or a properly licensed library voice. Test a short phrase, then record title, writers, instrumental source, performer, tool, voice model/licence, process date, and AI disclosure. Mark unknown fields [not provided] instead of inventing credits.

Evidence: clean vocal, two interpretations, instrumental, release note, and rights record.

Final Project - One Content Package

Required componentDeliverable
BriefAudience, objective, message, facts, constraints, channel
Visual identityCharacter bible and four consistent images
Main video45-60 second vertical story with narration and captions
AudioClean narration and music with clear usage rights
DocumentationPrompts, cost log, sources, permissions, and QA
ExportFinal file tested on a phone and editing device

Optional: your own-face presenter, lip sync, try-on, or a music-video section. Do not force face swap or cloning into a message that does not need it.

30-Day Study Plan

DaysFocusEvidence
1-3Fundamentals, brief, prompt A/B/CBrief and selected prompt
4-6Character bible and imagesMaster and four scenes
7-9Storyboard and videoThree shots and edit
10-11Script and voiceoverClean and mixed audio
12-15Own face and licensed swapPortraits, clips, consent, QA
16-18Story and rough cutSix beats and rough cut
19-20Animation and lip syncTest clip and timestamp QA
21-22Virtual try-onThree concepts and product report
23EfficiencyBudget and credit log
24-25Own voice/avatarSamples, consent, tests
26-27Song and music videoAudio and storyboard
28Vocal conversionTwo versions and release note
29AssemblyComplete main video
30QA and portfolioFinal, rubric, documentation

The 60-90 minute daily allocation is a suggestion. Extend it for hardware, credit, or provider constraints. Never skip consent or rights checks to meet the schedule.

Record the face/voice/asset owner, exact source files, permitted techniques, purpose and channels, script/context boundaries, term/territory, AI providers, access/storage/deletion, compensation/licence, name/date, and approved output version.

This worksheet is not a final commercial contract or legal advice. Do not promise that every copy already distributed online can be retrieved.

Rubric and Publishing Gate

Score each component from 0-4:

ComponentWeight
Brief and message15
Story and continuity15
Visual consistency20
Audio quality15
Editing and readability15
Rights, consent, transparency15
Documentation and efficiency5

A suggested internal target is 75/100, but stop publication when permission is unclear, output makes a false claim about a real person, product details materially change, or key facts remain unverified. Aesthetic quality cannot compensate for rights or identity failure.

Run three QA passes: content and facts; technical checks on headphones and phone speakers; then rights, consent, disclosure, and archive completeness.

Cross-Module Troubleshooting

SymptomLikely causeNext action
Output keeps driftingBroad or conflicting briefLock one objective, reference, and format
Identity changesInconsistent referencesReturn to the master; change one variable
Good images, unclear storyShots have no narrative functionRebuild beats before adding assets
Video breaks near the endAction is too complexShorten or simplify the shot
Text/logo is wrongGenerative typography is unstableAdd verified text manually
Voice sounds unnaturalScript is not spoken languageShorten sentences and test pronunciation
Lip sync driftsSource audio or timeline changedCheck audio and start point
Try-on alters the productModel invented detailsStop publication or use real photography
Credits disappear quicklyProduction started before decisionsReturn to the storyboard and attempt limit
Publishing rights are unclearLicence layers are incompleteFinish the rights record and seek professional review

You finish the course when you can show not only the final output, but also its brief, prompts, versions, licensed inputs, attempt log, QA evidence, and the reason behind each production decision.

Official sources and references

Use these sources to confirm current commands, capabilities, prices, and limits.

Was this guide helpful?

Tell us whether the steps worked or if something needs an update.

AI Content Creator: 15 Practical Modules | AI Blueprint Learning