Published onSeptember 9, 2026Renting a GPU to run an LLMllmgpuvllmrunpodinferenceServing a 70B model on rented GPU since I can't run that on my MacBook Pro, then breaking it on purpose and measuring recovery.
Published onSeptember 7, 2026Running your own LLMs (for beginners)llmgpuinferencevllmmachine-learningA working glossary of LLM inference workloads running on GPUs
Published onMay 9, 2025Plumber R reading AWS SageMaker custom attributesawssagemakerrplumberinferenceReading AWS SageMaker asynchronous inference `CustomAttributes` header in R using Plumber