Krea 2 Identity Edit
Identity-preserving instruction image editing on Krea 2
Apps started here then claimed by the 🐐 authors
Identity-preserving instruction image editing on Krea 2
Multi-view character sheet from one image (FLUX.2 LoRA)
Efficient native-resolution image generation and editing
Codec-native video & image understanding with Mage-VL 4B
Extend images into larger canvases with Krea 2 outpaint
Live interactive world rollout from an image
Layout-controlled comics with a draggable bbox canvas
Distilled LTX-2.3 identity video from a reference photo
Unified audio-text intelligence
Turn a 3D mesh into a parametric CadQuery program
Multilingual CPU-only ASR with a 1.58-bit BitNet decoder
Realtime VLM for image and video understanding
Bilingual EN/KO speech LM - transcribe, ask, and speak
Find when a described sound happens in a recording
Infographic generation & editing with SenseNova-U1 V3
Verbatim + intended transcripts with word-level timing
Studio-quality generative speech enhancement
Zero-shot TTS with explicit word-level prosody control
Aerial object detection - YOLO Models
Generative 4x video upscaling with LTX-2.3 IC-LoRA
Segment and track object instances in video with QueenVIS
Depth-controlled image generation with Krea-2 Turbo
Re-render a video from a new camera angle via IC-LoRA
Remove people & vehicles from video, keep the background
Parse document images into structured Markdown
1.06M-param pure-loop transformer with six effort levels
Style-guided image generation with Krea 2 Turbo
ClinFusion medical multimodal LLM for 2D images
Spatial physics video generation with MiniMax-H3 LoRA
FLUX.2 Klein 9B fine-tune for image generation and editing
Krea 2 Turbo HD text-to-image with enhanced VAE
Hierarchical parallel document parsing with a 1B VLM
Multilingual translation across 46 languages with MiLMMT
Steer a 3D camera path through any still image
Native size-sensitive matroid audit
Audio reasoning with evolving rubric rewards
Edit images or generate from Canny edges with NK2E on Krea 2
DINOv3 feature upsampling with ViT-Up
Tiny any-to-any multimodal GPT (text + image) prototype demo
LightOn multimodal document reranker with pointwise scoring
Cinematic product commercial style video LoRA for LTX-2.3
Multilingual PII detection & LLM safety moderation
Zero-shot CT findings with the Jolia 3D CT foundation model
Watch a 574K-param non-Transformer LM think
Play chess against a 32.5M-param search-free policy
Zero-shot probabilistic time-series forecasting
Occult philosophy chat LLM fine-tuned on Gemma 4 12B
Tiny 31.7M SLM that solves arithmetic expressions
Generate scene-linear HDR images with Krea 2 + LogC4 LoRA
Speech-driven talking-head video (Bernini-R S2V, Wan2.2)
Four voices in one 10M-param CPU model, with blending
Listwise document reranker with jina-reranker-v3.5
Photorealistic skin-texture LoRA for Krea 2 Turbo
MobileWan text-to-video generation (Qualcomm AI Research)
Compact multilingual ASR model (324M params)
Dense fine-grained image captioning with Qwen3-VL-4B
Proactive real-time commentary on audio-video streams
Predict 3D protein structures with OpenDDE
An RL plate that balances 21 objects in MuJoCo
Subject-driven text-to-video from reference images (Wan2.2)
Arabic speech to fully-diacritized text (tashkeel)
SAGE retrieves matching place images from a gallery
Convert molecular structure images into E-SMILES strings
AlienLM privacy layer for black-box LLMs
Compact 60M T5 translator for 15 languages
Unified AR model for image understanding & generation
Faithful x4 image super-resolution via FLUX.1-dev dual-LoRA
Zero-shot multilingual NER with GLiNER on LFM2.5-350M
Multi-speaker meeting transcription with diarization
Recognize text from WordArt / artistic scene text images
Relight exterior video clips with a light-direction ball
Image and video understanding with MOSS-VL multimodal model
Remove adverse weather effects (Histoformer, ECCV 2024)
Speak to a drone using an image, video, text, or voice.
Medical reasoning VLM based on Qwen3.6-27B
Agentic video understanding with active perception loop
Multimodal bio foundation model for molecules & proteins
Forecast 3D point trajectories from video and language
VLA model for scientific laboratory robotics
Distractor-free 2D enhancer for radiance field renderings
Fill-mask demo for a French medical ModernBERT
Speech-driven lip-sync on Wan 2.2 image-to-video
3D-centric world-spatial-action robot policy demo
Multi-shot narrated video with cross-shot memory
Multi-image instruction-guided image editing
Encode a context into a LoRA and chat with Qwen3-8B
Turn hardware ideas into structured JSON blueprints
Replicate camera motion from a reference video via IC-LoRA
Persian TTS demo using Ava-82M (Kokoro-based)
Define JSON tools and watch Lumma-0.6B generate tool calls
Multimodal reasoning VLM with Grug-style thinking
Detect watermarks, signatures, and artifacts in images
Proposal-only typed personal actions with a 35M-param model.
Multi-shot cinematic text-to-video generation
Generate market-grounded business ideas from images with MBA
Multilingual translation across 46 languages with MiLMMT
Rebuild any 3D mesh as a clean artist-style mesh
Video verification & temporal grounding with VideoSearch-R1
Hyperbolic vision-language zero-shot classification
Multilingual manga speech-bubble OCR (JA/ZH/EN)
4-step RL text-to-image with MeanFlowNFT on SD3.5-Medium
EN subtitle cues to Taiwan Traditional Chinese (0.6B model).
Predict robot action chunks from an image + instruction
Generate electron micrographs from text prompts
Forecast Valence & Arousal state change from posts
Graph-native LLM scientific reasoning with graph viz
Speaker-conditioned TTS with emotion & energy control
MrFlow training-free diffusion acceleration demo
Unified multimodal video generation and editing (5B)
Multi-agent interleaved text-image generation pipeline
P2R fine-grained visual reasoning with Qwen3-VL
VLM-guided tree-search keyframe extraction from videos
Human-object interaction video from image+audio+text
Facial affect estimation with uncertainty via rectified flow
Classify time series in-context with TimEE foundation model
Novel camera viewpoint from a video via depth-warp IC-LoRA
Ultrasound image understanding VLM with MoE architecture
Training-free typographic attack defense for CLIP
Reconstruct video from event-camera voxel grids via LongE2V
4-class lung ultrasound video classifier with attention
Detect AI-generated audio-visual content (DAV-Det)
Spatial reasoning VLM for 3D relations and perspective
Synthesize contrast-enhanced breast MRI from pre-contrast
Counterfactual robot futures from a world-action model
Multilingual translation across 46 languages with MiLMMT
Deobfuscate and detoxify Korean text with KOTOX
Sequence-only protein-ligand binding scorer (LULA-1.1)