Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Self-hosting LLMs on Google Cloud with GPUs using vLLM
Learn how to deploy the Gemma 2 model on Google Cloud GPUs using vLLM, covering setup, optimization, and cost‑effective self‑hosting.
Don’t want to be depedent on paid APIs for using LLMs? Then self-hosting is the way to go.
However, due to the “large” aspect of LLMs it’s extremely difficult to self-host them.
Everybody talks about self-hosting LLMs but no one actually does it.
vLLM is a framework that optimises the hosting of such LLMs 👉 https://github.com/vllm-project/vllm
I made a simple tutorial on how to deploy your own GEMMA 2 model on Google Cloud 👉 https://github.com/ThomasVrancken/info9023-mlops/tree/main/demos/04_vllm
Deploys Gemma 2 LLM using vLLM for accelerated inference on GKE GPUs.
vLLM: High-throughput, memory-efficient LLM serving engine using PagedAttention.
Compose Email
Loading recent emails...