Images and voices with your own keys: how media generation works in Personarium
BYOK at Personarium isn't just for text. Images, voice, and optional video follow exactly the same principle: your key, your provider, no Personarium server in between.
Images: one click, no second setup
By default, images run on the same OpenRouter key you already set up for text chat — no extra step needed. Want more model choice, you can optionally set up fal.ai as a second provider with its own key. Reference images keep your character's look consistent across generations, and the WebP variants for different display sizes are computed locally on your device (Rust image/libvips) — your generated image isn't sent anywhere again just to be resized.
Read-aloud with your own TTS key
For read-aloud (TTS), you connect with your own key at Inworld or ElevenLabs and assign each character its own voice. Generated audio is compressed as Opus and cached locally — if a character says the same line twice, you don't pay or wait twice for the same audio file.
Video is a feature flag with a built-in warning
Video generation is optional and runs on the same fal.ai key as additional image models (Kling, Veo, WAN, and similar) — image first, then image-to-video, so your character stays consistent. Runs take minutes and continue as a background job in a queue. Because a video costs 10 to 100 times more than a single image, Personarium mandatorily shows an estimated cost before every click — your key, but no surprise on the provider's bill.
What that means for you
Personarium takes no markup on any of these three — images, voice, or video; the membership doesn't touch these prices, they're an entirely separate bill with your chosen provider. The legally required safeguards (the age-18-or-older field, the prompt blocklists) apply identically across all media types, not just text.
Images, voice, video — the same rule as text chat: your key, your provider, no markup from us.