Define GPU functions in Python. Run them in the cloud with one decorator.
Documentation β’ Examples β’ Discord
pip install runpod # or: uv add runpod
rp login # authenticate onceRequires Python 3.10+. Installing the package also installs the rp CLI.
import runpod
from runpod import App, Model, NetworkVolume, Secret
app = App("inference")
models = NetworkVolume("models", size=100)
llama = Model("meta-llama/Llama-3.1-8B-Instruct")
# an autoscaling job queue on cloud H100s: weights pre-cached,
# dependencies vendored at deploy time, scale-to-zero when idle
@app.queue(
gpu="H100",
workers=(0, 3),
dependencies=["vllm"],
mounts={"/runpod-volume": models},
model=llama,
env={"HF_TOKEN": Secret("hf-token")},
)
def chat(prompt: str):
import vllm
llm = vllm.LLM(model=str(llama.path)) # weights already on disk
return llm.generate(prompt)
# one ephemeral pod per call: provisions, runs to completion, terminates
@app.task(gpu="H100", gpu_count=2, mounts={"/models": models})
def finetune(steps: int = 1000):
...
return {"loss": final_loss}
@runpod.local_entrypoint
def main():
print(chat.remote("why is the sky blue?")) # blocks for the result
job = finetune.spawn(steps=500) # fire and forget -> Jobrp flash dev main.py # live dev session: edit, re-run, logs stream back
rp flash deploy # deploy production endpointsPull requests and issues are welcome β see the contributing guide to get started.
git clone https://github.com/runpod/runpod-python.git
cd runpod-python
make setup
make testThis python package can also be used to create a serverless worker that can be deployed to Runpod as a custom endpoint API.
Create a python script in your project that contains your model definition and the Runpod worker start code. Run this python code as your default container start command:
# my_worker.py
import runpod
def is_even(job):
job_input = job["input"]
the_number = job_input["number"]
if not isinstance(the_number, int):
return {"error": "Silly human, you need to pass an integer."}
if the_number % 2 == 0:
return True
return False
runpod.serverless.start({"handler": is_even})Make sure that this file is ran when your container starts. This can be accomplished by calling it in the docker command when you set up a template at console.runpod.io/serverless/user/templates or by setting it as the default command in your Dockerfile.
See our blog post for creating a basic Serverless API, or view the details docs for more information.
You can also test your worker locally before deploying it to Runpod. This is useful for debugging and testing.
python my_worker.py --rp_serve_api.
βββ docs # Documentation
βββ examples # Examples
βββ runpod # Package source code
β βββ api # rest api v2 wrapper
β βββ cli # Command Line Interface Functions
β βββ endpoint # Language library - Endpoints
β βββ serverless # SDK - Serverless Worker
βββ tests # Package tests