-
New Jul 30, 2026
Under the Hood: Serving Kimi K3
DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a...
-
New Jul 23, 2026
Outperforming Fable 5 at half the price: meet model synthesis, a new server-side tool on DigitalOcean Inference Engine
Anyone building with AI runs into the same tradeoff: how to get the most intelligence per dollar, the right model at the right cost for each task. DigitalOcean Infer...
-
New Jul 21, 2026
Upcoming GPU Pricing Updates
Effective August 1st, 2026, we will be updating prices on select GPUs. This change reflects strong demand for advanced GPU capacity and helps us expand reliable access to high-performance compute fo...
-
New Jul 9, 2026
Scale Faster with Managed Weaviate: Now in Public Preview on DigitalOcean
Production Weaviate in minutes, managed by DigitalOcean. Starting at $20/month. Vector databases have become a core piece of the AI application stack. Whether you’re building retrieval-augment...
-
New Jul 2, 2026
Built for Mass Scale: Hard-Won Lessons from Teams Running High Volume Inference Workloads in Production
Moving AI from a flashy demo to a high-volume production environment is a transition filled with hidden technical debt and infrastructure challenges. There’s a difference between calling the OpenAI AP...
-
New Jul 1, 2026
DigitalOcean Evaluations: Production Model and Router Testing for the Inference Stack
Choosing the right model or inference router for production means more than reading a leaderboard. It means validating any model or routing configuration on your own data using your prompts and yo...
-
New Jun 25, 2026
Run Codex in the cloud – DigitalOcean for Codex is now available
As your agents are working on more complex, long-running work, they need a clean, persistent environment to keep running. Setting up a persistent remote machine by hand means creating a cloud se...
-
New Jun 17, 2026
Server-Side Tools Are Now Available for DigitalOcean Inference Engine
AI applications and agents are only as capable as the tools, data, and systems they can access. With Server-Side Tools, now in Public Preview for DigitalOcean Inference Engine, a model can call out to...
-
New Jun 10, 2026
The Inference Alpha: Maximizing Frontier Models on AMD
At DigitalOcean, we’re committed to providing high-performance infrastructure for the next generation of AI, which is why we’ve been focused on hosting frontier Large Language Models (LLMs) on fr...
-
New Jun 9, 2026
What We Learned Hiring 33 Engineers in Two Weeks
Earlier this year, we needed to hire a cohort of engineers in Seattle, fast. We had a product launching at our marquee conference, Deploy, a hard deadline, and a clear picture of what the work would a...
-
New Jun 4, 2026
Model Evaluations: Prove Your Routing Policy Actually Works
Most teams running inference at scale do not fail because they cannot find a “good” model. They fail because they ship a routing policy that looks fine in a playground, but drifts the moment it se...
-
New Jun 3, 2026
The Team Behind Deploy: Shipping AI, the DigitalOcean Way
Deploy 2026 came and went, and we’re still buzzing. For one day at Convene 100 Stockton in San Francisco, developers, startup founders, customers, and partners filled the room to talk about a shared c...
-
New Jun 3, 2026
Powering the Inference Era: Inside the DigitalOcean Data & Learning Layer
Building an AI-native application requires a data layer that can do two things at once: handle the structured, transactional queries your application runs on, and understand meaning well enough to po...
-
New Jun 2, 2026
Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era
The growth of generative AI isn’t driven solely by AI companies with proprietary models. Open-source AI is reshaping the developer ecosystem, fueled by a growing community of builders. But what do...
-
New Jun 1, 2026
The Inference Tax: How Prefix-Aware Routing Eliminates the Hidden Cost of LLMs at Scale
Introduction Inference demand is growing fast, and it’s only accelerating. By 2030, inference is expected to account for the majority of AI compute globally. But scaling inference i...
-
New Jun 1, 2026
DigitalOcean Serverless Inference: A Deep Dive
The Problem: Inference Gets Hard at Scale If you’ve shipped an AI feature to production, you already know: the hard part isn’t making a model resp...
-
New May 29, 2026
AI Disruptors: How the Next Generation of Business is Being Built
Getting your hands on a capable AI model is the easy part now. Every team can reach the same frontier models through an API, so a strong model is not what sets a product apart. What separates a wo...
-
New May 28, 2026
OpenCode Now Supports DigitalOcean Inference Router for Intelligent Model Routing
Coding agents today have a massive spending problem. Every request, whether you’re designing system architecture or writing a single-line docstring, often gets routed to the same expensive frontier mo...
-
New May 27, 2026
Scalable, Cost-Efficient AI: Introducing Unified Batch Inference on DigitalOcean
At Deploy 2026, we introduced the DigitalOcean AI-Native Cloud, built for the inference era. Batch Inference on the DigitalOcean Inference Engine enables high-volume asynchronous workloads. As develop...
-
New May 22, 2026
Request-Based Autoscaling Is Now Generally Available on App Platform
Traffic doesn’t spike on a schedule. A product launch, a viral moment, or a flash sale can send request volume through the roof in seconds, long before your CPU metrics catch up. That gap is where pe...
-
New May 20, 2026
How We Built DigitalOcean Inference Router
Most teams building on LLMs today make a single model decision and apply it uniformly across every request. They reach for a frontier model not because every task demands it, but because building th...
-
New May 13, 2026
Your Model Doesn't Matter. Your Infrastructure Does.
Everyone calling an LLM API has access to the same models. So what actually sets technical teams apart? It’s everything around the model like the routing logic, the live data pipelines, and the abili...
-
New May 4, 2026
Powering the Inference Era: Inside the DigitalOcean AI-Native Cloud
I’ve spent the last fifteen years building cloud services: early days of AWS building S3 and EBS, helping launch Oracle Cloud Infrastructure from inception, and now building the agentic cloud at Di...
-
New Apr 28, 2026
Introducing DigitalOcean AI-Native Cloud for Production AI Workloads
The AI industry has a compounding bottleneck, and it isn’t the models. It’s inference. What used to be a single model call has become a system of continuous interaction. Applications now orchestra...
-
New Apr 28, 2026
How we built the most performant DeepSeek V3.2, MiniMax-M2.5 and Qwen 3.5 397B on DigitalOcean Serverless Inference
Today at Deploy, we are announcing the general availability of DeepSeek V3.2, MiniMax-M2.5, and Qwen 3.5 397B on DigitalOcean Serverless Inference. On DeepSeek V3.2 and Qwen 3.5 397B, we deliver #1 ou...
-
New Apr 25, 2026
DigitalOcean Dedicated Inference: A Technical Deep Dive
Getting a model to answer 10 inference requests concurrently is tricky but simple enough; getting it to handle 2,000 engineers hitting a coding assistant with long contexts, all day, without ru...
-
New Apr 23, 2026
Beyond the Abyss Project Poseidon’s Quest for Zero-Downtime Reliability
In large-scale cloud environments, unpredictable hypervisor crashes carry real operational cost. While traditional reactive monitoring that relies on static thresholds and post-hoc alerts were on...
-
New Apr 23, 2026
From Incident Counting to SLIs: How DigitalOcean Rethought Availability
Our journey to truly understand our customer experience began with a hard look at our internal availability numbers at the start of 2025. We saw something uncomfortable: the numbers didn’t match ou...
-
New Apr 22, 2026
The LLM Inference Trilemma: Throughput, Latency, Cost
We know how to scale traditional web services: throw a load balancer in front of stateless microservices and horizontally scale your CPU instances as traffic grows. Large Language Models break this pl...
-
New Apr 21, 2026
Mastering the 600B+ Frontier: Optimizing Large Model Deployments on the Inference Cloud
We have moved past the point where a 70GB model was considered “heavy.” With the rise of models like DeepSeek-V3, the GLM series, and other massive Mixture-of-Experts (MoE) architectur...
-
New Apr 17, 2026
The Inference Cloud Memory Layer: A Technical Dive into DigitalOcean Managed Databases
As AI moves from experimental chat interfaces to production-grade agents, the need for a foundational memory layer to transform these AI-powered tasks into stateful models is apparent. The absence o...
-
New Apr 15, 2026
Load Balancing and Scaling LLM Serving
Load balancing for LLMs is fundamentally different from load balancing for traditional services like web servers, APIs, or databases. Prompt caching is the reason. Prompt caching typically cuts inp...
-
New Apr 13, 2026
Building a Robust Documentation Agent with DigitalOcean Gradient AI Platform
At DigitalOcean, documentation has always been a priority. Developers come to our docs to get unstuck, and the faster they find what they need, the better. Traditional docs pages work, but they re...
-
New Apr 7, 2026
Advanced Prompt Caching at Scale
Introduction Prompt caching is the process of reusing already computed KV states across inference requests in order to save money and reduce latency. Within a single replica, m...