Private models
You aren't limited to the public models on Replicate: you can deploy your own custom models using Cog, our open-source tool for packaging machine learning models.
Unlike public models, most private models (with the exception of fast booting fine-tunes) run on dedicated hardware so you don't have to share a queue with anyone else. This means you pay for all the time instances of the model are online: the time they spend setting up; the time they spend idle, waiting for requests; and the time they spend active, processing your requests. If you get a ton of traffic, we automatically scale up and down to handle the demand.
For fast booting fine-tunes you'll only be billed for the time the model is active and processing your requests, so you won't pay for idle time like with other private models. Fast booting fine-tunes are labeled as such in the model's version list.
Hardware pricing
| Hardware | Price | GPU | CPU | GPU RAM | RAM |
|---|---|---|---|---|---|
| CPU (Small) cpu-small | $0.000025/sec $0.09/hr | - | 1x | - | 2GB |
| CPU cpu | $0.000100/sec $0.36/hr | - | 4x | - | 8GB |
| Nvidia A100 (80GB) GPU gpu-a100-large | $0.001400/sec $5.04/hr | 1x | 10x | 80GB | 144GB |
| 2x Nvidia A100 (80GB) GPU gpu-a100-large-2x | $0.002800/sec $10.08/hr | 2x | 20x | 160GB | 288GB |
| Nvidia H100 GPU gpu-h100 | $0.001525/sec $5.49/hr | 1x | 13x | 80GB | 144GB |
| Nvidia L40S GPU gpu-l40s | $0.000975/sec $3.51/hr | 1x | 10x | 48GB | 65GB |
| 2x Nvidia L40S GPU gpu-l40s-2x | $0.001950/sec $7.02/hr | 2x | 20x | 96GB | 144GB |
| Nvidia T4 GPU gpu-t4 | $0.000225/sec $0.81/hr | 1x | 4x | 16GB | 16GB |
| Additional hardware | |||||
| 4x Nvidia A100 (80GB) GPU gpu-a100-large-4x | $0.005600/sec $20.16/hr | Additional Multi-GPU A100 capacity is available with committed spend contracts. | |||
| 8x Nvidia A100 (80GB) GPU gpu-a100-large-8x | $0.011200/sec $40.32/hr | Additional Multi-GPU A100 capacity is available with committed spend contracts. | |||
| 2x Nvidia H100 GPU gpu-h100-2x | $0.003050/sec $10.98/hr | Additional Multi-GPU H100 capacity is available with committed spend contracts. | |||
| 4x Nvidia H100 GPU gpu-h100-4x | $0.006100/sec $21.96/hr | Additional Multi-GPU H100 capacity is available with committed spend contracts. | |||
| 8x Nvidia H100 GPU gpu-h100-8x | $0.012200/sec $43.92/hr | Additional Multi-GPU H100 capacity is available with committed spend contracts. | |||
| Nvidia H200 GPU gpu-h200 | $0.001525/sec $5.49/hr | H200 capacity is available with committed spend contracts. | |||
| 2x Nvidia H200 GPU gpu-h200-2x | $0.003050/sec $10.98/hr | Additional Multi-GPU H200 capacity is available with committed spend contracts. | |||
| 4x Nvidia H200 GPU gpu-h200-4x | $0.006100/sec $21.96/hr | Additional Multi-GPU H200 capacity is available with committed spend contracts. | |||
| 8x Nvidia H200 GPU gpu-h200-8x | $0.012200/sec $43.92/hr | Additional Multi-GPU H200 capacity is available with committed spend contracts. | |||
| 4x Nvidia L40S GPU gpu-l40s-4x | $0.003900/sec $14.04/hr | Additional Multi-GPU L40S capacity is available with committed spend contracts. | |||
| 8x Nvidia L40S GPU gpu-l40s-8x | $0.007800/sec $28.08/hr | Additional Multi-GPU L40S capacity is available with committed spend contracts. | |||