A 100% free & open-source AI Content Automation Tool that writes scripts, generates voiceovers, creates videos, and uploads them automatically — hands-free YouTube growth powered by AI.
-
Updated
Sep 2, 2026 - Python
A 100% free & open-source AI Content Automation Tool that writes scripts, generates voiceovers, creates videos, and uploads them automatically — hands-free YouTube growth powered by AI.
新一代 AI 专业字幕软件,剪映字幕、elevenlabs 语音转文本的最佳本地版平替之一,也是加强版。中英转录识别准确率超过 97%,词语音频对齐率 98%,说话人分割与识别准确率 96%,带有最先进的 ASR 开源模型,100+字幕动画(支持导出透明背景的字幕动画)。说话人识别、专业字幕编辑器、命令行工具、Skill,达芬奇字幕插件(含字幕动画插件),PR 字幕插件,本地转录、远程转录、文稿匹配、智能拆行、AI校正、AI 智能热词、翻译、双语字幕、专业字幕编辑器、字幕合成、自定义大模型 API
[CVPR 2020] Meshed-Memory Transformer for Image Captioning
A neural network to generate captions for an image using CNN and RNN with BEAM Search.
[CVPR 2019] Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions
[ICLR 2026] An official implementation of "CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning"
Official Repository of OmniCaptioner
[CVPR 2021] Scan2Cap: Context-aware Dense Captioning in RGB-D Scans
[T-PAMI 2024] & [CVPR 2023] Vote2Cap-DETR; A set-to-set perspective towards 3D Dense Captioning; State-of-the-Art 3D Dense Captioning methods
ShapeGPT: 3D Shape Generation with A Unified Multi-modal Language Model, a unified and user-friendly shape-language model
CVPR 2018 - Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
Scalable annotation pipeline for action-aglined fine-grained instruciton for Visual-language-Action model
Computer Vision Playground ⚡️
[ECCV2022] D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding
Privacy first powerful, browser-based tool for tagging, captioning, cropping and managing training datasets for Stable Diffusion's LoRA trainings. No install required.
Implemented 3 different architectures to tackle the Image Caption problem, i.e, Merged Encoder-Decoder - Bahdanau Attention - Transformers
PyTorch library for Visual-Semantic tasks
Computer Vision: Generate captions that describe the contents of images using PyTorch
Cross-platform desktop client for AI.Opensubtitles.com featuring automatic transcription and translation of media files with FFmpeg integration and responsive UI.
To ease the driver to identify the Traffic Signs and also for the efficient working of Self-Driving Cars.
To associate your repository with the caption-generation topic, visit your repo's landing page and select "manage topics."