Inference Provider

VERIFIED
13,725 monthly requests

AI & ML interests

Text-to-image, image editing, text-to-video, speech recognition, text-to-speech, audio, inference providers, fine-tuning

Recent Activity

lucataco  updated a model 3 days ago
replicate/vllm-flash-attn3
lucataco  updated a model 3 days ago
replicate/yoso
lucataco  updated a model 3 days ago
replicate/triton_kernels
View all activity

Organization Card

Run AI with an API

Replicate lets developers run, fine-tune, and deploy open models with a production-ready API. On Hugging Face, you can use Replicate as an Inference Provider for popular models across image generation, image editing, video generation, speech recognition, and text-to-speech.

Featured models: Run with Replicate

All Replicate-powered models: Browse the full list

Read the integration docs: Replicate on Hugging Face Inference Providers

Examples for every task: Replicate Inference Provider Examples

Why use Replicate on Hugging Face?

Use Replicate infrastructure through Hugging Face's standard Inference Providers interface. You can call image, video, speech, and audio models with the same InferenceClient, your existing HF_TOKEN, and a provider switch instead of wiring up a separate integration path.

Get started in 30 seconds

Create a token at huggingface.co/settings/tokens with the "Make calls to Inference Providers" permission, then generate an image:

Python

pip install huggingface_hub pillow
export HF_TOKEN=hf_...
import os
from huggingface_hub import InferenceClient

client = InferenceClient(
    provider="replicate",
    api_key=os.environ["HF_TOKEN"],
)

image = client.text_to_image(
    "A cinematic photo of an astronaut riding a horse",
    model="Tongyi-MAI/Z-Image-Turbo",
)

image.save("replicate-astronaut.png")

JavaScript

npm install @huggingface/inference
export HF_TOKEN=hf_...
import { InferenceClient } from "@huggingface/inference";
import { writeFile } from "node:fs/promises";

const client = new InferenceClient(process.env.HF_TOKEN);

const image = await client.textToImage({
  provider: "replicate",
  model: "Tongyi-MAI/Z-Image-Turbo",
  inputs: "A cinematic photo of an astronaut riding a horse",
});

const ext = image.type.split("/")[1];
await writeFile(`replicate-astronaut.${ext}`, Buffer.from(await image.arrayBuffer()));

For cURL and task-specific examples, see the Replicate provider docs.

Billing

You have two options:

  • Bill through Hugging Face (default): requests are charged to your Hugging Face account at Replicate's rates with no markup, and your monthly HF credits apply. No Replicate account needed.
  • Use your own Replicate key: add your Replicate API token in your Inference Providers settings. Your code stays the same, and usage is billed directly by Replicate.

Team and Enterprise orgs can bill usage to the organization with bill_to="your-org" (Python) or { billTo: "your-org" } (JavaScript). See Pricing and Billing for details.

Go further on Replicate

Need fine-tuning, deployments, or thousands more models? Visit replicate.com and the Replicate docs.