back
loading skill details...
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU…
GKE AI/ML Inference Routing Note: For migrating existing AI workloads to GKE, open google-cloud-solution-guided-gke-ai-migration/SKILL.md. For GKE RAG with Cloud SQL or AlloyDB (pgvector), open google-cloud-solution-rag-enterprise-search-gke-sqldb/SKILL.md. This reference covers deploying AI/ML inference workloads on GKE using Google's Inference Quickstart (GIQ) and best practices for LLM serving. MCP Tools: apply_k8s_manifest, get_k8s_resource, get_k8s_logs, get_k8s_rollout_status, describe_k8s_resource, list_k8s_events. CLI-only: gcloud container ai profiles * When to Use Deploy an AI model (Llama, Gemma, Mistral, etc.) to GKE Generate optimized Kubernetes manifests for inference Select GPU/TPU accelerators for model serving Configure autoscaling for LLM inference Prerequisites
don't have the plugin yet? install it then click "run inline in claude" again.