IP-Adapter (Image Prompt Adapter) is a lightweight adapter module for pre-trained text-to-image diffusion models that enables image-conditioned generation by injecting reference image features via a decoupled cross-attention mechanism. Introduced by Tencent AI Lab in 2023, it allows users to supply a reference image alongside a text prompt to control style, subject identity, or composition without fine-tuning the base diffusion model. The adapter architecture inserts parallel cross-attention layers that process image embeddings from a pre-trained image encoder such as CLIP, keeping base model weights frozen.
Content
- IP-Adapter was introduced by Ye et al. from Tencent AI Lab in a 2023 paper titled “IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.” The central problem it solves is image prompt compatibility: standard diffusion models accept only text prompts, but many practical applications require controlling generation with a reference image—for example, preserving a subject’s face, copying an artistic style, or anchoring composition from a sketch or photograph.
- The technical approach is elegant: rather than replacing or fine-tuning cross-attention layers, IP-Adapter adds a parallel cross-attention branch specifically for image tokens. Image features are extracted from a frozen CLIP image encoder, projected into the text embedding space via a trainable linear layer, and then attended to by the new parallel branch. The text cross-attention branch remains unmodified, preserving the base model’s text-following behaviour. During inference, the two attention streams are combined with a controllable weight, enabling smooth blending between text-driven and image-driven generation.
- A key practical advantage is parameter efficiency: only the projection layer and parallel attention weights are trained, totalling roughly 22 million parameters for SD v1.5 models. Training on large-scale paired image-caption datasets (such as LAION) takes orders of magnitude less compute than full model training, and the resulting adapter is base-model-agnostic—a single adapter can be applied to any fine-tuned checkpoint derived from the same base architecture.
- IP-Adapter-FaceID extends the base design with face-specific identity embeddings, enabling consistent face generation across diverse scenes, poses, and styles without identity drift. IP-Adapter-Plus variants incorporate multiple reference images and support higher-resolution feature injection. These specialisations demonstrate the adapter pattern’s extensibility for domain-specific conditioning.
- The adapter has become a widely adopted tool in creative and commercial pipelines, used for product visualisation (placing a product image in a generated scene), character consistency across illustrated narratives, and virtual try-on applications. Its integration with ComfyUI and AUTOMATIC1111 makes it accessible to non-programmers. The pattern it establishes—decoupled cross-attention for modal conditioning—has influenced subsequent adapter architectures in the diffusion model ecosystem.