AI Model Hub
Every model in one place — image, video, and text. Run any of them on CNAPS with no setup and no code.
AI Models
- Depth Anything AnnotatorMonocular depth estimation (Depth Anything V1 Large).
- Gender RecognitionClassify pedestrian gender (male/female).
- Adult Content DetectionClassify image safety (safe / nsfw).
- Object ClassificationImageNet-1k classification (ResNet).
- Image ColorizationAuto-colorize B&W photos (DISCO). 10-30s.
- Real-World Blur RemovalDiffusion-based deblur (highest quality). 1-3 min.
- Motion Blur Removal (MSSNet)Balanced motion-blur removal (MSSNet). 5-15s.
- Image DenoiserDenoise photos (SwinIR). 5-15s.
- FireRed-1.1Natural-language image edit with strong identity consistency. 1-3 min.
- QWEN-2511Edits images from a natural-language instruction. Higher quality than QWEN 2509 but heavier (1–3 min).
- ControlNet XL CannyGuided SDXL image generation from a Canny edge map + text prompt. Sharper structural adherence than Union. Outputs up to 1024×1024.
- ControlNet XL UnionGuided SDXL image generation from a control image + text prompt. One model, 6 modes: canny, depth, openpose, hed, normal, segment. Outputs up to 1024×1024.
- BLIP (Image Description)Auto image captioning (BLIP Large). Fast.
- Gemma4-E2BGemma 4 E2B compact VLM, 128K context, bf16.
- Gemma 4 31BGemma 4 31B Dense flagship VLM, 256K context, thinking mode.
- ViLT (Image Q&A)Visual Question Answering (ViLT, VQAv2).
- Qwen3.6-35B-MoEQwen 3.6 35B/A3B MoE VLM, 256K context. fast inference.
- LatentDiffusion (Object Removal)Mask-based object removal / inpainting. 10-30s.
- DeepSeekOCRMulti-language + handwriting OCR. 30-60s.
- GLM-OCRDocument OCR with layout/table/formula. 10-40s.
- PaddleOCRKorean-optimized printed OCR. Fast.
- DETR (Face Detector)Face detection (DETR ResNet-50 fine-tuned). Fast.
- DETR (License Plate Detector)License-plate detection (DETR ResNet-50). Fast.
- DETRGeneral object detection on 91 COCO classes (DETR).
- YOLO-SFast object detection, 91 COCO classes (YOLOS-Small).
- YOLO-S (Fashion Item Detection)Fashion item detection — 46 apparel classes (YOLOS).
- RF-DETR MediumReal-time object detection on 80 COCO classes (RF-DETR Medium).
- RF-DETR KeypointReal-time human pose estimation predicting COCO 17 keypoints per person, rendered as an OpenPose-format skeleton map (RF-DETR Keypoint Preview).
- QWEN-InpaintPrompt-driven mask inpaint — fills the masked area from your prompt, keeps the rest.
- QWEN-LayeredDecomposes an image into editable RGBA layers (foreground objects, background, …). Pair with List-Extract → QWEN-Image-Edit → List-Inject → Layer-Compose to edit a single layer and recompose.
- JPEG Quality RestorationRemove JPEG compression artifacts (SwinIR). 5-15s.
- PiSA-SRHighest-quality image upscaler (2x/4x). ~50s.
- SMFANet+Fast lightweight image upscaler (2x/3x/4x). 0.1-3s.
- Swin2SRBalanced image upscaler (2x/4x) — Swin Transformer V2. 10-30s.
- SwinIRImage upscaler with widest scale range (2x/3x/4x/8x). 10-30s.
- BiRefNet (Subject)Cleanly extracts a single salient subject with high-resolution edges. ~1-3s.
- RF-DETR Seg Medium (Scene)Real-time instance segmentation on 80 COCO classes (RF-DETR Seg Medium).
- SAM2 (Scene)Segment every object in an image (no labels needed). 10-30s.
- SAM3.1 (Scene)Language-prompted segmentation — SAM 3.1 Object Multiplex with improved accuracy over SAM3. 10-30s/class.
- SAM3 (Scene)Language-prompted segmentation (e.g. "person", "red car"). 10-30s/class.
- Cosmos3 NanoNVIDIA world model for text- and image-to-video, with optional synced audio.
- Helios BaseOn-device text-to-video, up to 60s @ 24fps. 3-10 min.
- Helios DistilledFast distilled text-to-video, fixed 640x384. 1-5 min.
- MiniMax-H3 Ref2VAMiniMax-H3 Ref2VA (omni-reference) generates a 5-15s 24fps video with native stereo audio from a prompt plus a reference image and a reference clip, that clip's own soundtrack included. Video and its soundtrack come out of one denoising loop, so lip movement and sound land in sync. ~144GB bf16 weights — offloaded component-by-component, minutes-scale per clip.
- Wan2.2 TI2V 5BWan2.2 dense 5B unified text+image-to-video. Text only → 5s video from prompt; text + image → image as the first frame of the video.
- Subject Follow (9:16)Click one person and get a 9:16 vertical video that follows only them, audio intact.
- SAM3.1 (Video)Language-prompted video segmentation and tracking — SAM 3.1 Object Multiplex. Up to 16 objects tracked jointly. 1-3 min per 200 frames.
- FlashVSR v1.1One-step diffusion 4x video super-resolution (FlashVSR Tiny pipeline). Input resolution up to 720×720.
- SeedVR2 3BVideo upscale to 720p–4K (SeedVR2 3B). 2-5 min (longer for 4K).
- SeedVR2 7BMax-quality video upscale to 720p–4K (SeedVR2 7B). 5-15 min (longer for 4K).
- SparkVSRDiffusion video upscale (2x/3x/4x) with optional ref image. Input resolution up to 480×480.
- Speech Recognition (Clips)Recognize speech in each clip of a video list and convert it to timed transcripts.
- Speech Recognition (Video)Recognize speech in a video and convert it to text.
External Models
- Claude.aiAnthropic's Claude LLM. Generates text from a prompt — strong at reasoning, coding, and long-document analysis.
- Nano BananaGoogle Gemini-based image gen/edit with multimodal reasoning. Needs API key.
- VEOGoogle VEO 3.1 video gen (up to 4K, with native audio). Needs API key.
- GeminiGoogle Gemini cloud LLM. Strong reasoning + multimodal. Needs API key.
- GPT ImageOpenAI GPT-Image text-to-image / edit. Needs API key.
- Sora 2OpenAI Sora 2 text-to-video / img2video. Needs API key.
- ChatGPTOpenAI GPT cloud LLM. General-purpose reasoning. Needs API key.
- OmniGoogle Gemini Omni: text/image/video in, video out, with conversational editing. Which input ports apply depends on Task — Text to Video: Text only. Image to Video: + Image 1. Reference to Video: + Image 1–3. Edit: + Video (aspect ratio is ignored). Unused ports are dropped even if wired. Needs API key.
- Text to Speech (Gemini)Google Gemini TTS: turn text into a spoken voiceover (Korean/English). Needs API key.
- Video AnalysisGemini video analysis. Feed a video and get a transcript, timecodes, or highlight analysis as text. Needs API key.
IO Nodes
- Font LoaderLoad a font file (.ttf/.otf) as workflow input.
- Gallery ViewerDisplay an image list as a thumbnail gallery.
- Image LoaderLoad an image file as workflow input.
- Image ViewerDisplay the final image output.
- Sound LoaderLoad a sound file as workflow input.
- Sound OutputPlay the final sound output.
- Text InputProvide a text string as workflow input.
- Text ViewerDisplay text output (OCR, captions, LLM).
- Video LoaderLoad a video file as workflow input.
- Video ViewerDisplay the final video output.
Tools
- Canny Edge AnnotatorClassical Canny edge detection.
- Audio FinishNormalize a video's loudness to a broadcast target (2-pass loudnorm + peak limiter).
- Image CompareSide-by-side image viewer (2-4 inputs).
- Text CompareSide-by-side text viewer (2-4 inputs).
- Video CompareSide-by-side video viewer with synced playback (2-4 inputs).
- Color Pick by ClassExtract a segmentation class by name.
- Color Pick by RGB ValueExtract segmentation regions by exact RGB color.
- Manual Color PickerInteractively pick colors from a segmentation map.
- Image Crop by ClassCrop detected objects by class name.
- Image Crop by CoordCrop a region from explicit bounding-box coordinates.
- Image Gate by TextPass or block an image based on text match.
- Image SelectorPick first non-empty image from inputs (fallback).
- Image Masker by ClassBuild a B&W mask from segmentation/detection by class name.
- Image Masker by CoordBuild a B&W mask from explicit bbox coordinates.
- Image Masker by RegionBuild a B&W mask from line-separated polygons or bounding boxes.
- Manuel Image MaskerHand-draw a B&W mask with brush tools.
- Object PickerInteractive — click segmented objects to build a mask.
- Text ConcatenateConcatenate up to 4 text inputs into one output.
- Text DecoratorWrap text with a template (placeholder substitution).
- Layer ComposeAlpha-composite a layer stack into a single image.
- List ExtractExtract one image from an image list by index.
- List InjectReplace one image in a list by index with a new image.
- Region ExtractExtract region polygons or bounding boxes from OCR region data.
- Color GrayscaleConvert color image to grayscale. Instant.
- Color InverseInvert pixel colors (photo negative). Instant.
- Image Blur (Fast)Fast blur via downsample-then-upsample.
- Draw TextOverlay text on an image with font/size/color control.
- Image Blur (Standard)Standard Gaussian blur with iteration control.
- Image AddAdd two images pixel-wise (lighten / double-exposure).
- Image Conditional ResizeAuto-resize image to the next standard resolution tier.
- Image MultiplyMultiply two images pixel-wise (cutout / multiply blend).
- Image Resize To MatchResize image to match another image's dimensions.
- Image ResizeResize image to a specific width/height.
- Lens BlurCircular disc-kernel blur for camera-bokeh look.
- Manual CropCrop or fill a drawn region (rect/circle/polygon).
- Selective BlurBlur only a drawn region (rect/circle/polygon).
- Image Blur (Simple)Single-slider blur (1-10). Fastest blur option.
- Video Frame ExtractExtract a single frame from a video as an image.
- Video ReassembleReassemble an image list into a video.
- Video SplitSplit a video into a list of image frames.
- Video TrimCut a source video into a list of clips, one per highlight segment.
- Video CaptionBurn each clip's highlight text onto it as a caption overlay.
- Video Clean SpanTrim each clip to its single dominant continuous shot (remove mid-clip frame jumps).
- Video ConcatConcatenate a video list into a single video.
- Video Constraint MapMeasure a source video's editing constraints for a shorts planner.
- Video EDL PlannerChoose the exact cut and framing inside each candidate take.
- Video EDLTurn an edit plan into the segments the shorts chain renders.
- Hook Teaser (Cold Open)Prepend a cold-open teaser of each clip's hook moment to the front.
- Video List ExtractExtract one video from a video list by index.
- Video Narrate (Qwen3)Prepend a spoken voiceover cold-open to every clip in a batch (Qwen3).
- Narration Mux (Cold Open)Prepend a voiceover cold-open (frozen first frame + narration) to a clip.
- Video Reframe (9:16)Reframe each clip in a video list to 9:16 vertical (1080x1920).
- Video SFXMix caption-synced, emotion-matched sound effects into each clip.
- Remove Dead ZoneRemove silent gaps from each clip in a video list to tighten pacing.
- Slice Transcript to ClipsSlice a whole-video transcript onto each clip.
- Snap Segments to SpeechSnap segment boundaries to speech word boundaries.
- Video SpeedRetime each clip to a punchier tempo (pitch-preserving).
- Video Subtitle (Spoken)Burn word-synced (karaoke) subtitles onto each clip from its transcript.