2026 Edition · Our own software · local or cloud APIs

The agentic studio for AI video.

One idea in. A finished, on-brand vertical reel out — storyboard, footage, voiceover, music, presenter and publishing, on a pipeline you can run yourself or drive with an agent.

9:16Native vertical · 16:9 · 1:1
1Idea → finished reel
6+Local engines, one pipeline
~70Agent tools over MCP
01 · Making short-form video in 2026

Great videos are easy to imagine.
They're brutal to actually make.

01

Every reel is six tools.

A script writer, a footage source, a voiceover app, a music library, a video editor, a scheduler. Six tabs for one 20-second clip.

02

Stock footage looks like everyone else.

The same drone shots, the same gradients. Nothing on screen is yours, and nothing is on-brand.

03

Voice and music are a licensing maze.

A consistent narrator costs a subscription. Royalty-free music that does not sound royalty-free costs another.

04

On-brand means redoing it every time.

The logo, the voice, the company line — pasted back in by hand on every single export.

05

Cloud video bills by the second.

Per-clip pricing makes volume expensive, and your raw footage lives on someone else’s server.

06

Five platforms, five exports.

YouTube, TikTok, Instagram, Facebook, LinkedIn — five logins, five captions, five aspect ratios.

02 · One idea, nine stages, one reel

Type the idea. The studio does the rest.

Each stage is its own model and its own editable step — run the whole thing in one call, or stop and tune any stage before the next.

01

Scenario Wizard

Ollama · local LLM

Type a plain-language idea. The local model returns an ordered storyboard of 3–12 cinematic shots — framing, motion, mood, all editable.

02

Narration

Ollama

AI-writes a spoken line per scene, sized to each shot’s duration and matched to a tone you set — warm, punchy, authoritative.

03

Shots

ComfyUI

Each scene renders as text-to-video, or image-to-video when the shot has a reference frame. Re-roll any single shot without touching the rest.

04

Voice

XTTS

Per-scene voiceover or one continuous track. Use a built-in speaker, or your brand’s cloned voice from a short sample.

05

Music

ACE-Step · Suno

A scored soundtrack from style tags, BPM and length — generated locally with ACE-Step, or via Suno with your own API key.

06

Presenter

Wan 2.2 S2V

Optional: turn any scene into a lip-synced talking presenter from a single face image. Fast or high-quality pass.

07

Cover

SDXL

A branded intro still generated to open the reel — your logo and palette baked in.

08

Assemble

Remotion · ffmpeg

Shots, voiceover and soundtrack stitched into one finished 9:16 reel, captions and overlays in place.

09

Publish

Scheduler

Draft a caption, schedule a slot, and post to YouTube, TikTok, Instagram, Facebook and LinkedIn from one screen.

▶
make_reel — the whole run in a single call.Scenario → narration → render every shot + music + speech → assemble → finished reel. One tool, hands-free.
03 · By the numbers

Built to make video at volume.

3–12Cinematic shots per storyboard, from one idea
6+Local engines — video, image, voice, music, LLM, lip-sync
5Publish targets from a single draft
~70MCP tools — drive the whole studio from any AI client
3Aspect ratios — 9:16, 16:9, 1:1, set per reel
0Per-clip cost when you run it on your own hardware
04 · Two ways to render

Render on your GPU, or on cloud APIs you bring.

It's one self-hosted studio — you decide where each stage runs. Own a GPU? Render locally for free. Prefer the cloud? Drop in your own provider keys. Mix them per capability.

ON YOUR GPU

Local render

Your GPU. Your models. Zero per-clip cost.

The whole pipeline runs on your own hardware. Open models do the work, nothing leaves your network, and the marginal cost of another reel is electricity.

  • ComfyUI for text-to-video & image-to-video
  • Ollama for the storyboard, narration & captions
  • XTTS voice synthesis + brand-voice cloning
  • ACE-Step music, Wan 2.2 S2V presenter, SDXL covers
  • Remotion + ffmpeg assembly, a worker queue you control
  • Full privacy — footage and scripts never leave your box
See the engines →
YOUR API KEYS

Cloud APIs

No GPU? Bring your own provider keys.

No hardware to spare? Drop in your own cloud API keys and the studio routes that stage to a hosted provider instead — you pay the provider directly, nothing runs through us.

  • Your own keys — e.g. Kling for video, Suno for music
  • Latest hosted models, no drivers or downloads
  • Set a key per capability; leave the rest local
  • You pay the provider directly — no markup, no middleman
  • Same storyboard → reel → publish workflow
  • Mix and match: local where it is cheap, cloud where it is not
See the engine matrix →
05 · The engine room

Open models locally. Your API keys in the cloud.

Every capability is a swappable slot. Run it on open weights you control, or point it at a cloud provider with your own key — per capability, your call.

CapabilityLocal · your GPUCloud API · your key
VideoComfyUI — text-to-video & image-to-videoKling (your API key)
Image · CoverSDXL via ComfyUIYour image API key
Script · LLMOllama (any pulled model)Your LLM API key
VoiceXTTS — synthesis + voice cloningYour TTS API key
MusicACE-StepSuno (your API key)
Presenter · Lip-syncWan 2.2 S2VYour lip-sync API key
AssemblyRemotion + ffmpegLocal (both routes)
06 · Brand kit

Every reel comes out on-brand.

Set it once per project. Logo, voice, presenter and company details merge into every render — no copy-paste on export.

Logo

Dropped onto covers and overlays so every reel is unmistakably yours.

Cloned voice

A short sample becomes a reusable narrator — the same voice across every video.

Presenter face

One face image powers lip-synced talking-head scenes on demand.

Company details

Name, phone, website and email merge into captions and end cards automatically.

Brand voice card

A written tone description steers the storyboard and narration writers.

Tag catalogs

Editable libraries of cinematic scene concepts and music styles, reused across projects.

07 · You stay in the director's chair

Generated, but not on autopilot.

The wizard gives you a first cut in one pass. From there every shot, line and beat is yours to nudge — and only what you change re-renders.

  • Edit any shot · rewrite the line, swap the reference image, change motion — then re-render just that scene
  • Reorder & trim · drag shots, adjust durations; narration resizes to fit
  • Rewrite voiceover · regenerate a single shot’s spoken line without touching the video
  • Overlays · text, lower-thirds and brand graphics composited at assembly
  • Aspect per reel · 9:16 for shorts, 16:9 for YouTube, 1:1 for feeds
  • Jobs you can watch · every render is a queued job — poll, wait, or cancel from the dashboard or an agent
Storyboard · “Atelier launch”9:16 · 6 SHOTS
01Macro: thread pulled through linenRENDERED
02Hands shaping clay on the wheelRENDERED
03Presenter: “Made by people, not machines.”PRESENTER
04Product turn on a warm-lit pedestalRENDERING
05Logo cover · champagne on charcoalQUEUED
Music · ACE-Step · cinematic, warm · 92 BPM
08 · Agentic by design

Drive the whole studio from your AI client.

LuwiStudios ships a Model Context Protocol server. Claude Desktop, Cursor, n8n or your own agent can run the entire pipeline — in plain language.

MCP ClientCLAUDE · CURSOR · N8N
→
LuwiStudios MCP~70 TOOLS
→
Render workersCOMFYUI · OLLAMA · XTTS
create_projectcreate_scenariowrite_narrationgenerate_shotgenerate_speechgenerate_musicgenerate_avatarassemble_reelmake_reelcreate_post+ ~60 more
09 · Why now

The economics of video just flipped.

Format

Short-form won.

Vertical 9:16 is the default unit of attention across every major platform. The brands that post daily are the brands that get seen.

Models

Open weights caught up.

Wan, SDXL, XTTS and ACE-Step now produce broadcast-credible output on a single GPU box — quality that used to mean a cloud bill.

Cost

Per-second pricing breaks at scale.

Cloud video is fine for a handful of clips. At volume, self-hosting turns marginal cost into electricity — and keeps your footage in-house.

Frequently asked questions

Answers to the most common questions about LuwiStudios.

Our own agentic studio for AI video — end-to-end software we build in-house. You give it an idea; it writes a storyboard, renders each shot, generates voiceover and music, optionally adds a lip-synced presenter, and assembles a finished vertical reel ready to publish. Every stage renders on your own GPU, or through cloud APIs you bring your own keys for — your choice, per capability.

You already have the ideas.

LuwiStudios turns them into finished video — on your hardware or ours, by hand or by agent. Tell us what you want to make.

Our own software · ComfyUI · Ollama · XTTS · ACE-Step · Wan 2.2 · SDXL — or bring your own cloud API keys (Kling · Suno · …)