AI Video Generator from Text and Images
Create videos from a prompt or a source frame. See the exact price before you start

About the service
AI video generator from text and images
Create short videos from a text prompt, animate photos, and control motion with a first or last frame. Capabilities depend on the model: selected models support reference images and videos, audio, vertical 9:16 output, and resolutions up to 1080p. The form always shows the current settings.
How to create an AI video
- Choose a model and mode: text, first frame, last frame, first and last frames, or reference media.
- Describe the scene, camera movement, action, lighting, and visual style.
- Sign in and select the duration, format, resolution, and audio when supported by the model.
- Check the exact price and start the task. Download the completed video from the result page.
Processing time depends on the model, duration, resolution, and current load. This tool generates a new video rather than editing an existing one, so review the result before publishing it.
Current models
| Model | Suitable workflows | Payment |
|---|---|---|
| Grok Imagine Video 1.5 | Text or first-frame video, 4–15 seconds, with audio | Paid balance |
| MomentFlow v5.2 | Text, frames, and reference media, 4–30 seconds | Paid balance |
| MomentFlow v5 Fast | Faster generation from text, frames, and references | Paid balance |
| WildClips v4.3 | First, last, or both frames, 5–15 seconds, with audio | Paid balance |
| WildClips v4.1 | Frame-controlled generation, 5–15 seconds | Paid balance |
| Framy NSFW v1.7 | First-frame video, 4–30 seconds, adults only | Bonus or paid balance |
| Framy NSFW v1.5 | First-frame video, 5–30 seconds, adults only | Bonus or paid balance |
Adult models require age confirmation. All generations must follow the service rules and applicable law. The catalog may change; the form retrieves current models and their available settings from the API.
Photo animation, social clips, and text-to-video
- Image to video: upload a starting frame and describe the subject or camera movement.
- Text to video: describe the scene, characters, action, lighting, and composition.
- Reels and Shorts: choose 9:16 output on a model that supports it.
- First and last frames: control the start and end of a transition.
- Reference media: use images or videos as guidance on compatible models.
Price and daily bonus
No subscription is required: you pay for the selected generation, and the exact price appears before launch. The bonus balance can be restored up to $0.60 once per UTC day. Recent paying users receive it automatically; other users claim it on the Billing page.
For video generation, the bonus balance currently works only with Framy NSFW v1.5 and v1.7. All other models require paid funds. The bonus does not provide unlimited generations or access to every model, and it must cover the full task price. See the pricing page for details.
The Problembo interface has no third-party ad placements. Problembo does not add its own visible watermark to downloaded videos.
Frequently asked questions
Is registration required?
You can browse the model catalog and calculate a price without signing in. An account is required to upload files and start a generation.
Which files can I upload?
Frames and reference images support JPG, JPEG, PNG, and WebP files up to approximately 15 MB. Compatible models also accept MP4, MOV, and WebM reference videos up to 100 MB.
Can I use the result commercially?
Commercial use is allowed for generations paid while the paid balance is positive, subject to third-party rights and the selected model's terms. Bonus-funded generations should not be used commercially.
What's new
- Reduced video generation prices for MomentFlow v5.2 across all available resolutions, both with and without a video reference.
- Also added support for 1080p resolution.
- Added xAI's new video model, Grok Imagine Video 1.5. It creates videos from text or images up to 15 seconds long at 480p or 720p.
- The model generates audio alongside the video, including speech, music, sound effects, and ambience, while synchronizing lip movements with dialogue.
- Users particularly praise its speed, faithful preservation of the source image, and short-scene quality. Complex motion and voice generation can be inconsistent, so the model works best for concise clips with a clear prompt.
- Added MomentFlow v5.2. The new version generates videos up to 30 seconds instead of 15, handles longer narratives and scene transitions better, and maintains stronger visual consistency. Up to 30 images and 10 videos can be used as references.
- Reduced MomentFlow v5 Fast generation prices across all available resolutions, both with and without a video reference.
- WildClips v4.0 is now available with native stereo audio: the model handles dialogue voiceovers especially well, synchronizes speech with on-screen action, and adds environmental sounds.
- Eleven languages are supported reliably: Russian, English, Chinese, Arabic, French, German, Italian, Japanese, Korean, Portuguese, and Spanish.
- Videos can be generated at up to 2K resolution and up to 15 seconds long.
- Text, first and last frames, images, video, and audio can be used as references (soon).
- MomentFlow v5 and MomentFlow v5 Fast now support attaching a video file as a reference.
- A reference video helps preserve motion and style from the source clip, for example when you want to continue a video.
Framy has been updated to version 1.5, and this is no cosmetic patch — it's a major engine upgrade.
What's New
Image Quality
The latent space has been rebuilt with an updated VAE trained on higher-quality data. In practice this means: hair finally looks like hair, text is legible, and fine facial details no longer dissolve into mush. Visual artifacts have been noticeably reduced.
Audio Out of the Box
Framy 1.5 generates synchronized audio in a single pass — dialogue, ambience, sound effects. An updated vocoder has improved speech intelligibility, and cross-modal alignment has reduced lip-sync and timing drift.
Motion and Consistency
The image-to-video mode has been significantly improved: fewer "frozen" frames, fewer spontaneous panning shots, and better preservation of the source image's visual integrity.
Native Vertical Video
For the first time, vertical generation up to 1080×1920 is supported — the model is trained on portrait data rather than cropping from landscape. Reels and Shorts can now be fed directly.
The text encoder has been scaled up 4×, making the model significantly better at understanding complex prompts describing camera angles, character movements, and scene composition.
Prompt Enhancer
Another major bonus is a powerful prompt enhancer. It helps turn a rough request into a more precise and expressive scene description, carefully refining motion, camera angles, dialogue, and atmosphere so the final video lands closer to the original idea on the first try.
In short, Framy 1.5 is one of those cases where the version number is being modest.