AI Image Generation
AI generates images from your conversation to bring chats to life. Describe a scene and see it visualized alongside your messages.
Learn how multimodal AI chat systems are transforming interactions by integrating text, images, voice, and more into seamless conversational experiences.
Multimodal is an easy word to put on a page. These are the specific things a Soulkyn thread can carry, each of them shipped.
AI generates images from your conversation to bring chats to life. Describe a scene and see it visualized alongside your messages.
Send an image into the chat and your kyn understands it — you can ask about a photo, and they can attach a picture to a reply of their own.
Not push-to-talk: the mic streams continuously with live voice-activity detection and a soft end-of-turn, captions re-transcribe as you speak, and the reply arrives sentence by sentence as it is synthesised. Mid-call you can change your kyn's language, your own, or force their original untranslated voice. Call time is unmetered.
Our in-house voice engine starts speaking in 0.07 seconds and generates 6× faster than realtime — 175K+ voices served on 4 dedicated voice clusters. Calls that never keep you waiting. Voices resolve automatically per kyn from gender and language, with a force-original-voice toggle for the untranslated version.
There is a mic button in the composer: record a voice message and it is transcribed into the chat.
Energy, Trust, Affection, Arousal, Mood and Pain track from 0 to 100%, update live and feed back into behaviour. That is what emotional context looks like when it is a number the system actually reads, not a claim on a landing page.
Two engines — a fast multilingual lane and a slower, higher-quality one for full songs. Pick a duration from one to four minutes or leave it on auto, toggle instrumental, let the lyrics be written from your kyn's personality or write them yourself with section tags, and add a style caption. Music can be generated straight from a chat message and attaches to it. 640K+ songs and clips have been rendered so far.
Any image in any gallery can become a 5 or 10 second clip, with sound if you choose the Video + Sound mode. Premium required, and generation runs in the background with a notification when it lands.
AI remembers conversation history including preferences and past topics. Conversations build context across sessions without repeating information.
Soulkyn runs two tiers of text model. Which one your kyn speaks with depends on your plan — nothing to configure, nothing to switch on.
Standard chats run on Gwem, our fast in-house model: quick, solid, great for everyday roleplay.
Deluxe chats run on GLM 5.3, an Opus-class model — the size class of the biggest frontier assistants. Deeper memory, sharper dialogue, characters that truly keep track of the story. And it's unlimited, because it runs on our own hardware.
Soulkyn self-hosts a 753B-class frontier model on private clusters — which is exactly why unlimited is possible here and nowhere else. Deluxe and Deluxe Backer text chat uses it automatically; it does not apply to voice calls.

Chat using text or voice with image sharing capabilities. AI processes all input types together for unified conversation context.
Explore Now
Define character traits including personality style and conversation preferences. Adjust voice settings and response patterns to match what you want.

Combine text voice and images in single conversations with context preserved across all formats. Group chats support multiple AI characters interacting together.
ExploreCommunicate through words and receive AI-generated images that visualize your descriptions, creating a richer conversation experience.
Speak naturally to your AI companion and receive intelligent responses that understand tone, accent, and conversational context.
Energy, Trust, Affection, Arousal, Mood and Pain track from 0 to 100%, update live and feed back into behaviour. That is what emotional context looks like when it is a number the system actually reads, not a claim on a landing page.
Share images with your AI companion and receive thoughtful responses based on visual content and context.
Enjoy conversations that remember your preferences, history, and context across all modalities for truly personalized interaction.
Switch effortlessly between text, voice, and visual interactions while maintaining consistent conversation context.
Multimodal AI processes text, images, voice, and other media formats. Instead of text-only responses, these systems understand context across different formats for more natural conversations.
Not push-to-talk: the mic streams continuously with live voice-activity detection and a soft end-of-turn, captions re-transcribe as you speak, and the reply arrives sentence by sentence as it is synthesised. Mid-call you can change your kyn's language, your own, or force their original untranslated voice. Call time is unmetered.
Energy, Trust, Affection, Arousal, Mood and Pain track from 0 to 100%, update live and feed back into behaviour. That is what emotional context looks like when it is a number the system actually reads, not a claim on a landing page.
Memory is per-message vector recall plus chained rolling summaries plus forced memories you write yourself, all over one store, with a new message recallable within a couple of minutes.
Concretely rather than vaguely: a Share Generation Data toggle sits in the image generation controls and decides whether your prompt and settings travel with a public image. Forced memories are scoped to a single thread. Novels and personas are private by default and publishing is a deliberate act — and for novels, reversible. Voice calls are unmetered and stay on the standard model; the Deluxe model applies to text chat only.
Engage in open and respectful experimentation with AI interactions in a secure environment.
Pages people read next.