Caption Engine

HOW IT WORKS

From real video frames to ready-to-post ideas

Caption Engine studies what appears in an uploaded clip before selecting three short captions and ten relevant, viral-friendly hashtags.

1. Upload one video

The generator accepts MP4, MOV, and WebM files up to 200 MB. A successful generation costs one prepaid credit. The uploaded file is used only for the requested operation and is removed after processing finishes.

2. Extract chronological frames

The backend uses FFmpeg to sample real frames across the video. Looking across the sequence helps the model distinguish a subject, action, scene, visual style, emotional tone, important objects, and other supported context. It is instructed not to invent identities it cannot recognize confidently.

3. Check useful public context

When relevant, the AI can research current public information connected to the visible topic, such as franchise terminology, fandom discussion, subject names, or discoverable phrases. Public context supplements the clip; it does not replace the visual evidence.

4. Evaluate multiple candidates

The engine develops at least ten caption candidates and twenty-five hashtag candidates internally. It scores combinations for content match, search relevance, hook strength, niche fit, trend evidence, specificity, and discovery balance.

5. Return one precise set

The final response contains exactly three captions of eight words or fewer and exactly ten unique hashtags. Generic reach tags are penalized, while specific and publicly supported terminology is preferred.

What “viral-friendly” means

It means the set is prepared for realistic discovery using relevance, current public evidence, searchable language, niche specificity, and controlled breadth. TikTok’s private recommendation system is not available to Caption Engine, so virality and reach are never guaranteed.

Generate from your video

Use the clip itself as the starting point for captions and hashtags.

Open Caption Engine