Today
Front-end, in one read.
The latest from the sources worth following, newest first.
DEV CommunityThe Giantsread at source
I Built an AI That Turns “I’m Bored” Into Real-World Side Quests 🌿
This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass . Your neighborhood has side quests. Go unlock one. 🌿 What I Built POV: You open your phone to check one notification. An hour later, you're watching a raccoon steal someone's lunch and wondering where your life went. We've all been there. 💀 I wanted to build something that uses AI to get people off their screens instead of giving them another reason to stare at one. Meet OUTLAND: TouchGrass , an AI-powered IRL side quest generator. You tell it your mood, energy level, available time, budget, social preference, and how chaotic you're feeling. It turns that information into a small adventure you can actually do outside. Think of those days when you want to go somewhere, do something, touch grass, become the main character for five minutes, but your brain has absolutely no plans. That's where OUTLAND comes in. The goal isn't to gamify every second of your life. No streak anxiety. No guilt for staying home when you need a rest. Just a little nudge toward doing something different. Demo 🌱 Live app: https://outland-touchgrass.onrender.com/ Open the app, choose your current state, and generate a quest. When location information is available, OUTLAND can discover real places and ground the quest in a location that passes its checks. During testing, it generated a quest called THE GREEN LUNG AUDIT at Leisure Valley Park. Instead of simply saying “go for a walk,” it framed visiting a green space as a quest with a time estimate, difficulty, and XP. It's a small twist, but that's the idea. You don't need to book a trip, spend a lot of money, or coordinate six people's calendars just to do something different. Give it a try when you're bored, restless, or tired of doing the same thing every weekend. Your next side quest might be closer than you think. Code 💻 Open-source repository: https://github.com/LovelySharma-dev/Outland---TouchGrass The code is open for you to explore, learn from, and contribute to. If you have ideas for better quests or ways to improve the experience, I'd love to hear them. How I Built It The stack is Node.js, Express, JavaScript, Gemma, and SerpApi . Here's how the pieces work together. 📍 SerpApi discovers real places I didn't want the AI inventing a park with a suspiciously convenient name. SerpApi helps discover actual places so the generation pipeline has real location data to work with. 🧠 Gemma generates the quest The app combines structured user preferences with relevant place information. Gemma then generates a quest intended to fit the user's time, budget, energy, and social preferences. 🛡️ The app checks the result A response isn't automatically a good quest just because an AI generated it. OUTLAND checks safety, location grounding, quality, and personalization. Weak results can be rejected, with a bounded retry path. 🧺 When AI has a bad day, there is a fallback Free inference can be unreliable. Sometimes the model produces a generic quest, and sometimes the inference route is unavailable. OUTLAND includes a curated fallback pool for those situations. Fallback quests are labelled honestly rather than pretending they came from the model. I'd rather show a useful fallback than serve AI-generated nonsense with confidence. 🧪 What I learned The biggest lesson was that getting AI to generate text is one thing. Getting it to generate something a person might genuinely want to do is another. A quest can sound creative but still be too vague, ignore the user's energy level, or fail to point to a useful real-world destination. Free inference adds another challenge because model availability and response quality aren't always consistent. So I had to build around the model instead of treating the prompt as magic. Quality checks, bounded retries, location validation, and fallback behavior became important parts of the application. The takeaway? A clever AI demo and a useful AI product are not the same thing. The model can suggest an adventure. The system still has to make sure that adventure makes sense. Why Does Open Innovation Matter? I wanted to explore an AI use case beyond another chatbot or productivity assistant. Gemma powers quest generation, while SerpApi helps connect generated ideas to real places. The surrounding application handles validation, personalization, and failure cases. Using an open-weight model makes the generation layer more open to experimentation. Developers can inspect the project, change prompts, improve validation, and explore alternative model setups. I'm not claiming the whole app runs locally or without an internet connection. Instead, this project explores how open AI building blocks can power a practical experience that encourages people to spend less time on their screens. For me, the exciting part is not just what the model can generate. It's what we can build around it. My Agent Session I built OUTLAND with AI-assisted development. I'd like to share the build process so people can see how the idea evolved into a working application. Agent session: Add your DevRelay session link here if you saved one. This section is optional, so remove it if you don't have a session to share. Prize Categories Remove this section if you're not entering a specific partner prize category. If you are, list only the categories that genuinely apply to your project. 🌿 Your Turn OUTLAND was built for Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass . If you try it, tell me what quest you get. Did it give you a reason to explore somewhere new, or did you immediately return to scrolling? Be honest. 😂 🌱 Try OUTLAND: https://outland-touchgrass.onrender.com/ 💻 Explore the code: https://github.com/LovelySharma-dev/Outland---TouchGrass 🔗 Connect with me on LinkedIn: https://www.linkedin.com/in/lovely-sharma-dev/ If you're into AI, open source, or building fun things that solve real problems, let's connect. Your neighborhood has side quests. Go unlock one. 🌿 hacktoberfest #devchallenge #hf26challenge #gemma #opensourceai #javascript
DEV CommunityThe Giantsread at source
I made a JSON to Excel tool that can't upload your data, even if it wanted to
I kept getting JSON from APIs that I needed to look at as a table or send to someone who only uses Excel. The obvious fix is one of the many online JSON-to-Excel converters. But most of them send what you paste to their server, and I didn't love doing that with real data. So I made my own: jsoncsvvisualizer.com . It does two things. You can paste JSON or CSV and see it as a table, then download it as .xlsx. And there's a JSON formatter that also fixes broken JSON and tells you where it was broken. It's React + Vite + Tailwind, SheetJS for the Excel file, and it's just static files on Render. There's no backend. A few parts were more interesting than I expected, so I'm writing them down. Letting the browser enforce "we don't upload your data" Every converter site says it doesn't keep your data. I wanted something stronger than a sentence in a privacy policy. The site sends this Content Security Policy header: default-src 'self'; script-src 'self'; connect-src 'self'; connect-src 'self' means the page can only make network requests to its own domain. So even if some dependency got compromised and tried to send your data somewhere, the browser would block it. You don't have to trust me, which I like. One thing that tripped me up: Vite's dev server needs inline scripts and websockets for hot reload, so the strict policy breaks npm run dev . I ended up with a relaxed policy for dev and the real one for vite preview , so if I break the CSP I find out locally instead of in production. Finding where JSON is broken People paste stuff like this: { name : ' EleStack ' , tools : [ ' Grid ' , ' JSON ' ,], active : True } The jsonrepair library fixes this fine. But I didn't want to quietly change someone's data without telling them, so the tool shows a warning like "input had errors, line 3 col 26", and clicking it jumps the cursor there. Getting that line number was harder than I thought. My first attempt read the position out of the JSON.parse error message. Turns out Chrome, Firefox and Safari all word that message differently, and Safari often doesn't include a position at all. So I gave up on parsing error messages and wrote a small strict JSON scanner, about 60 lines, that walks the text and stops at the first character that breaks the rules. Then it's just: const before = text . slice ( 0 , position ); const line = before . split ( ' \n ' ). length ; const column = position - before . lastIndexOf ( ' \n ' ); Same result in every browser. Data pasted twice This one I only found by testing: if you accidentally paste the same array twice, you get [...] [...] , which isn't valid JSON, and the repair step makes a mess of it. The fix was to split the input into separate top-level values first, by counting brackets (and ignoring brackets inside strings), then repair each piece and merge the arrays. for ( let i = 0 ; i text . length ; i ++ ) { const ch = text [ i ]; if ( inString ) { /* skip until the closing quote */ continue ; } if ( ch === ' " ' ) { inString = true ; } else if ( ch === ' { ' || ch === ' [ ' ) { if ( depth === 0 ) start = i ; depth ++ ; } else if ( ch === ' } ' || ch === ' ] ' ) { depth -- ; if ( depth === 0 ) chunks . push ( text . slice ( start , i + 1 )); } } Records that don't match API data is rarely uniform. One record has email , the next doesn't. If you take the columns from the first row, you lose fields. So the columns are every key that appears in any row, and missing values are just empty cells: const keys = new Set (); rows . forEach ( row => Object . keys ( row ). forEach ( k => keys . add ( k ))); Nested objects get shown as JSON text so nothing disappears in the Excel file. Things I got wrong I launched with the same page title on every page, which is pretty bad for search. And the navbar was too wide on phones, so every page scrolled sideways. I only noticed after it was live. Both are fixed now. If you want to try it jsoncsvvisualizer.com . It's free, with no account and no cookies. If you have a really broken JSON file, I'd like to know if it survives. I'm also deciding what to add next, probably CSV to JSON or a tree viewer. Let me know which you'd actually use.
DEV CommunityThe Giantsread at source
What to Send a Developer Before They Build Your Trading Bot
Most custom trading bot projects that go wrong don't fail in the code. They fail in the brief. The trader knows the strategy by feel, the developer guesses at the gaps, and the finished bot does something slightly different from what the trader meant. Here's the one-page spec I ask every client for before building an MT5 Expert Advisor. If you fill this in, any competent developer can quote accurately and build what you actually want. 1. Market and timeframe Which symbols (for example EURUSD, XAUUSD, US30)? Which timeframe are signals calculated on? Which broker and account type (hedging or netting)? 2. Entry rules Write each condition as a plain sentence: "Go long when the 20 EMA closes above the 50 EMA on the H1 chart." "Only if RSI(14) is below 70 at that close." Say whether the signal is checked at bar close or on every tick. This one detail causes more backtest-versus-live surprises than anything else. 3. Exit rules Stop loss: fixed pips, ATR-based, or below the last swing? Take profit: fixed, ATR-based, or none? Trailing stop and breakeven: when do they activate, and by how much? Time exit: close after a maximum number of bars? Opposite signal: does it close the trade, reverse it, or do nothing? 4. Position sizing and risk Fixed lot size, or a percentage of the account per trade? Maximum number of open trades at once, and per symbol? Daily and weekly maximum loss before the bot stops trading? 5. Filters Trading hours or sessions to allow or avoid. Maximum spread allowed for an entry. News avoidance, if any. 6. What "done" looks like Backtest period and the results you expect to roughly reproduce. Inputs you want to be able to change without touching the code. Whether you need alerts (Telegram, email, push) for entries and exits. A tip that saves money If you can, send screenshots of three or four past trades marked up with exactly where you would have entered and exited, and why. Developers can test their logic against those examples, and you'll catch misunderstandings before a single line is written. Why this matters A clear spec means a fixed, honest quote, a faster build, and a bot that behaves the way you trade on your best day. I build custom MT5 Expert Advisors with this process, plus a proper risk and execution layer in every build. If you've got your rules written down and want them automated, have a look at CustomTradingBots.com .
DEV CommunityThe Giantsread at source
Multi-Model Routing and SSE Keep-Alives for Agent Resilience
TL;DR: Provider outages break workflows: Recent Anthropic API disruptions (October 5-6, 2026) affecting Claude Opus 5.5, Mythos 5.1, and Fable 5.1 demonstrate the risk of single-provider reliance for long-running agent tasks. Preventing silent timeouts: MonkeysCode shipped a model-proxy fix on October 4 to keep Anthropic Server-Sent Event (SSE) streams alive during extended model "thinking" phases, bypassing aggressive load balancer idle timeouts. Multi-model routing: Workloads can route to Azure Foundry GPT-6, which was integrated on October 5 with cache and effort parity, ensuring tasks complete even when primary endpoints degrade. Managing cache trade-offs: The new Auto mode tab (shipped October 6) tracks cache ratios and warns on mixed frontier provider configurations to help developers balance resilience and API costs. Building autonomous AI coding agents requires managing a significant amount of state. When an agent is tasked with planning a feature, editing multiple files across a codebase, and running tests, it relies on a continuous, reliable stream of reasoning from a frontier model. However, the underlying infrastructure powering these models is distributed, complex, and subject to disruption. On October 5 and 6, 2026, Anthropic experienced incidents causing elevated error rates across their Claude Opus 5.5, Mythos 5.1, and Fable 5.1 endpoints. According to reports from Statusgator and Claude.com, this disruption impacted Claude AI, the API, Claude Code, and Claude Cowork before recovering within twenty minutes. For developers relying on a single provider, an API outage or rate spike mid-workflow often results in lost state, hanging requests, or a complete halt to development tasks. When an agent is halfway through a complex refactor, a sudden 502 Bad Gateway or a silent timeout is more than an inconvenience; it is a loss of valuable context and compute time. To protect developer productivity, AI coding platforms must be engineered with network resilience and multi-model failovers as foundational requirements. Here is a technical look at how MonkeysCode handles connection management and multi-model routing to prevent agentic workflows from stalling during provider disruptions. The Mechanics of SSE Keep-Alives During Model Thinking When you invoke Capuchin , MonkeysCode's built-in coding agent, you are initiating a stateful, multi-step process. The agent reads the local file system, establishes context, and opens a streaming connection to an LLM provider using Server-Sent Events (SSE). SSE is a lightweight, unidirectional protocol built on top of HTTP, ideal for streaming text tokens from an LLM. However, during complex refactors, frontier models enter a "thinking" or reasoning phase. In this phase, the provider's API might not emit text tokens for several seconds or even minutes as it evaluates the prompt and computes the optimal output trajectory. At the network level, this silence presents a significant problem. Most standard reverse proxies, cloud load balancers, and client-side HTTP libraries enforce strict idle timeouts. If no bytes are transmitted over the TCP connection within that window (often 30 to 60 seconds), the intermediary drops the connection. The client receives a sudden ECONNRESET or a silent hang, and the agent's context window is lost. To mitigate this, MonkeysCode shipped a model-proxy fix on October 4, 2026, specifically designed to keep Anthropic SSE streams alive while the model thinks. By injecting periodic, non-disruptive keep-alive pings (typically formatted as SSE comments like : keep-alive\n\n ) into the stream at the proxy layer, we ensure the connection remains active through the load balancer's idle timeout window. This prevents silent drops during long-running planning phases, ensuring the agent can successfully transition from thinking to emitting actionable code edits without the connection being severed by overzealous network infrastructure. Multi-Model Routing: Bypassing Endpoint Degradation Even with robust connection management, a complete endpoint outage requires a different strategy. When Anthropic's Opus 5.5 endpoints experienced elevated error rates on October 6, developers locked into a single provider were left waiting for resolution. Resilience requires redundancy. MonkeysCode supports multi-model configurations, allowing developers to route requests across different frontier engines, local models, or through a Bring Your Own Key (BYOK) setup. On October 5, 2026, we expanded this redundancy by integrating GPT-6 via Azure Foundry, complete with cache and effort parity alongside our existing Anthropic support. By maintaining parity in how context is cached and how reasoning effort is parameterized, developers can switch models without rewriting their system prompts or losing performance baselines. If a primary provider experiences a spike in 5xx errors, developers can immediately route the agent's next step to a secondary provider. Because the MonkeysCode editor and Agent Manager maintain the local state and file context, the transition happens cleanly. The new model reads the existing context and continues the task. Furthermore, to optimize this transition, we shipped an update to our context-kernel on October 5 that keeps GPT-6 and 5.6 system blocks separate for cache breakpoints. This architectural decision ensures that when you do route to a different model, the system prompt structure aligns with that specific provider's optimal caching strategy, reducing unnecessary token processing. Managing Cache Trade-Offs in Auto Mode Switching providers mid-task introduces a specific trade-off: prompt caching. Frontier models rely heavily on prompt caching to reduce latency and API costs when processing large codebases. When you switch from Anthropic to Azure Foundry, the new provider does not have your codebase in its cache. The first request to the new provider will incur a full context-window read, increasing latency and cost for that specific turn. To make these trade-offs visible, MonkeysCode introduced the Auto mode tab on October 6, 2026. This interface provides direct visibility into your cache ratios. Crucially, it warns developers when they configure mixed frontier providers in Auto mode. If you configure Auto mode to round-robin or fallback between Claude Opus 5.5 and GPT-6, you risk thrashing the cache on both providers. The Auto tab surfaces these metrics, showing exactly how much of your context is hitting the cache versus requiring a fresh read. We also updated our usage dashboards on October 5 to explicitly show cache-write tokens. This visibility allows senior engineers to make informed decisions: accept the cache-miss penalty for the sake of immediate failover during an outage, or pin to a single provider and wait for the disruption to resolve to maximize cache utilization. The platform does not force a specific path; it provides the telemetry and the routing controls necessary to adapt to real-world infrastructure conditions. Takeaway Agentic coding workflows are only as reliable as the infrastructure executing them. By implementing strict SSE keep-alives to survive long thinking phases, and providing transparent multi-model routing with cache parity, MonkeysCode ensures your development environment remains functional even when frontier providers degrade. You can review your caching metrics and configure your model routing today in the MonkeysCode editor or via the Agent Manager . For teams looking to standardize these resilient workflows across their organization, explore our team plans and documentation .
DEV CommunityThe Giantsread at source
Claude Haiku 5.5 vs Opus 5.5: Pricing & API Guide
By Beat API Team · Prices and documentation checked October 9, 2026 Start with Haiku 5.5 for bounded, verifiable tasks. Start with Opus 5.5 when the task requires sustained reasoning across files, tools, or conflicting evidence. For applications containing both kinds of work, evaluate a Haiku-first route with explicit acceptance checks and an Opus escalation path. The Claude Haiku 5.5 vs Opus 5.5 comparison turns on pricing, cache behavior, and the cost of completing your task: Haiku's official uncached input/output rates are 40 times lower for prompts up to 100,000 tokens, but only eight times lower above that boundary. Cache reads count toward that boundary. A small new message attached to a large cached conversation still belongs to the long-prompt tier. Identical context limits do not imply identical ability to complete a task—or identical bills. Measure cost per accepted result, including retries and escalation. A cheap first attempt can leave expensive rework. API migration includes request parameters and response parsing; changing the model ID alone is insufficient. Disclosure: We build BeatAPI, which lists both models. This article separates Anthropic's published specifications, BeatAPI's dated public retail listing, and our own arithmetic and implementation recommendations. We have not run a controlled Haiku-versus-Opus quality or latency benchmark for this article. Claude Haiku 5.5 vs Opus 5.5: what actually differs? Both models accept text and image inputs and return text. Both advertise a 1M-token context window and a normal maximum output of 128K tokens. Haiku 5.5 was released on October 7; Opus 5.5 on September 22. Haiku is positioned for classification, extraction, routing, and subagent work; Opus for long-running coding and knowledge work. These are model specifications, rather than guarantees for every gateway, account, or integration. See the Haiku overview and Opus overview . That distinction suggests a useful first split: Workload Initial candidate Acceptance check Escalation trigger Classify a support ticket Haiku Allowed label, explicit intent, known exceptions Multiple intents or policy ambiguity Extract a field from a document Haiku Value exists in a cited source span Conflicting records or missing evidence Find likely files for a bug Haiku Existing paths and relevant symbols Scout cannot justify the shortlist Fix a multi-file regression Opus Reproduction, regression tests, diff review Review or test failure Reconcile conflicting documents Opus Traceable claims and resolved contradictions Evidence remains incomplete Write a summary for a recurring report Haiku for extraction; evaluate Opus for synthesis Source coverage and figure checks Cross-document inference or high rework rate This is a proposed routing policy. A simple task with unusual domain rules can still be difficult, while a large document with one exact lookup may be easy. Start with representative examples from your application before assigning a whole category to either model. Haiku 5.5 pricing vs Opus: input, output, and cache costs The following are Anthropic standard API prices in USD per million tokens , excluding Fast mode, Batch discounts, geography modifiers, and separately billed tools. Token category Haiku: prompt ≤100K Haiku: prompt >100K Opus: standard New input $0.10 $0.50 $4.00 Output $0.50 $2.50 $20.00 Cache read $0.01 $0.05 $0.20 5-minute cache write $0.125 $0.625 $5.00 1-hour cache write $0.20 $1.00 $8.00 For uncached input and output, Opus/Haiku is 40× below or at 100K and 8× above it. For cache reads alone, those ratios become 20× and 4× . A workload mixing these buckets has its own ratio. Rates and the counting rule come from Anthropic's pricing documentation . The boundary applies to the whole request , rather than a marginal surcharge on tokens after 100K. Prompt length includes new input, cache reads, and cache writes. Output does not determine prompt length, but its rate changes when the prompt crosses the boundary. Earlier requests retain their original prices. This matters when a conversation grows gradually. One additional tool result can move the next request into a different tier even if most of its context is cached. BeatAPI's separately verified retail listing On October 9, the anonymous BeatAPI pricing endpoint listed these rates for the exact IDs claude-haiku-5-5 and claude-opus-5-5 : Token category BeatAPI Haiku: standard BeatAPI Haiku: long_context BeatAPI Opus New input $0.05 $0.25 $2.00 Output $0.25 $1.25 $10.00 Cache read $0.005 $0.025 $0.10 5-minute cache write $0.0625 $0.3125 $2.50 1-hour cache write $0.10 $0.50 $4.00 These published rates are half the corresponding Anthropic standard rates. The listing verifies a public price book; it does not independently establish output equivalence, uptime, end-to-end latency, or support for every Anthropic feature. The endpoint exposes both Haiku tiers; the 100K rule described above is independently documented by Anthropic. Boundary settlement through BeatAPI has not been tested here. Recheck current pricing before budgeting. For an implementation check, the Claude Haiku 5.5 API page and Claude Opus 5.5 API page on BeatAPI bring the model IDs, current token rates, and request examples together. Use those examples to test one representative task before adopting the routing policy below. Five workloads you can calculate before calling either model For disjoint token buckets, a token-only estimate is: cost = (new_input × input_rate + cache_read × read_rate + cache_write_5m × write_5m_rate + cache_write_1h × write_1h_rate + output × output_rate) / 1,000,000 Do not add cached tokens to both new input and cache read. For the examples below, output means the full billed output token count , including reasoning where applicable, rather than just the visible answer. The token counts are assumed identical between models to isolate pricing; real model runs will usually differ. Assumed request Anthropic Haiku Anthropic Opus BeatAPI Haiku BeatAPI Opus 2K new input + 500 output $0.00045 $0.018 $0.000225 $0.009 90K cache read + 10K new input + 2K output $0.0029 $0.098 $0.00145 $0.049 190K cache read + 10K new input + 2K output $0.0195 $0.118 $0.00975 $0.059 100,000 new input + 2K output $0.011 $0.440 $0.0055 $0.220 100,001 new input + 2K output $0.0550005 $0.440004 $0.02750025 $0.220002 These are illustrative calculations, not observed invoices or benchmark results. Cache-hit examples exclude the earlier cost of creating the cache. BeatAPI columns apply its published tier rates with Anthropic's documented boundary as the budgeting assumption. Two consequences are easy to overlook: First, the cache-heavy 200K example makes Opus about 6.05× as expensive, rather than 40×. Most of the prompt is cheap cache-read traffic, while Haiku has crossed into its higher tier. Second, at this output length, increasing an uncached Haiku prompt from 100,000 to 100,001 tokens increases the estimated request cost roughly fivefold. Opus's estimate changes only by the additional input token. This is why token counting belongs before routing, especially near the boundary. For the small-request example, one million requests would cost $450 on Anthropic Haiku or $18,000 on Anthropic Opus. BeatAPI's listed rates imply $225 or $9,000. Those totals assume one attempt per request, unchanged token usage, no tools, and no other charges. They are planning scenarios, not promised savings. When is compaction worth paying for? Suppose the next ten requests each read 190K cached tokens, add 10K new tokens, and generate 2K output tokens. At Anthropic rates, Haiku costs 10 × $0.0195 = $0.195 . If a validated summary lets each request read 90K cached tokens instead, the ten requests cost 10 × $0.0029 = $0.029 . That leaves $0.166 for creating the summary and writing its replacement cache before the token-only saving disappears. At BeatAPI's listed rates, the corresponding allowance is $0.083. For example, writing a 90K replacement cache at Haiku's five-minute standard rate costs $0.01125 at Anthropic rates. The cost of generating the summary is additional. You also need to account for cache expiration and misses. The important qualification is semantic: a cheaper summary that loses the one detail needed later can increase retries or produce a wrong answer. Preserve source references and measure downstream acceptance. Compaction is an optimization to evaluate, rather than a universal instruction to shorten everything. What the published benchmarks can—and cannot—tell you Anthropic reports 39.2% for Haiku 5.5 and 66.4% for Opus 5.5 on Terminal-Bench 4.0 , and 46.4% versus 54.4% on FrontierCode 1.1 Main . The reported GDPval-AA v2.1 scores are 1620 and 1846 , respectively. Sources: the Haiku announcement and Opus announcement . These published results help choose what to test first. They do not provide your application's success probability. The Opus announcement specifies xhigh effort for Terminal-Bench and generally max effort for other reported results, and describes safeguard fallbacks in some evaluations. These figures are not a controlled comparison at equal effort, equal cost, or through BeatAPI. Avoid turning them into claims about a specific production route. Some apparently comparable visual scores also use different tool conditions or subsets. A result with tools and a result without tools should not be presented as a clean model-only ranking. Likewise, Elo scores are not percentages: a higher GDPval score does not imply a proportional increase in accepted tasks. A useful evaluation separates three questions: Capability: Can the model complete the task under the allowed tools and context? Economics: What does an accepted result cost with the actual request sequence? Product experience: Does it meet the latency and correction burden users can tolerate? Keep these measurements separate. The least expensive accepted answer may still arrive too late for an interactive workflow. A Haiku-first route needs a rejection policy A practical candidate architecture is: Task → explicit complexity rules ├─ known complex task → Opus → acceptance check └─ bounded task → Haiku → deterministic checks ├─ accepted → return └─ unresolved → Opus with source evidence Use observable evidence for the gate. For extraction, require a source span and validate the requested field. For code, run the reproduction and regression tests. For classification, check the allowed categories and evaluate difficult labeled examples. A model saying “I am confident” is not an acceptance test. Do not let Opus see only Haiku's conclusion when escalating a disputed result. Pass the original source, the relevant excerpt, and the failed check; identify Haiku's draft as an unverified candidate. Otherwise, escalation can amplify the first model's mistake. The escalation break-even formula Let H be average Haiku cost per incoming task, Oe the Opus cost when escalated, Od the cost of sending that task directly to Opus, V validation overhead, and p the escalation fraction. A one-attempt Haiku route costs: expected_cost = H + V + p × Oe It is cheaper than direct Opus when: p If escalation and direct Opus cost the same and validation is free, this simplifies to p . With the small-request assumptions above, H = $0.00045 and Oe = Od = $0.018 . If 20% escalate, the result is $0.00405 per incoming task , 77.5% below sending every task directly to Opus. At BeatAPI's listed rates it is $0.002025 versus $0.009, with the same relative saving under identical assumptions. This is a sensitivity example, not a measured 20% escalation rate . If Opus needs to reread more context or correct a damaged draft, Oe may exceed Od . Add all failed attempts, validator calls, tools, and review time before claiming a production saving. Also measure false acceptance : how often an incorrect Haiku result bypasses escalation. Lower escalation is not an improvement if the gate silently returns more wrong answers. Migration traps that affect the comparison Haiku 5.5's migration guide documents changes beyond the model ID: Old assumption Adjustment Token counts measured on Haiku 4.5 still apply Recount using the new model; the tokenizer can produce approximately 30% more tokens for the same text Fixed thinking.budget_tokens controls reasoning Use adaptive thinking and output_config.effort temperature=0 makes extraction predictable Remove sampling parameters; use explicit requirements and validation content[0] is always the answer Collect blocks whose type is text A tiny max_tokens cap is sufficient Reasoning can consume the cap before visible text appears Assistant prefill enforces JSON Replace prefill with a supported output mechanism and validate the result Opus has its own differences: adaptive thinking cannot be disabled and forced tool use returns an error. A shared router should maintain model-specific capability settings , rather than passing every Haiku request option unchanged to Opus. See the Opus overview . A minimal two-model request BeatAPI's public gateway contract exposes POST /v1/messages . This example uses that route and asks both models the same bounded question. It is an integration example; authenticated execution and model-specific parameter forwarding have not been tested for this article. Confirm them with a small canary before using the code in production. export BEATAPI_API_KEY = 'YOUR_API_KEY' # Save the Python example below as compare.py, then: python3 compare.py import json import os import time import urllib.request key = os . environ [ " BEATAPI_API_KEY " ] prompt = """ Classify this ticket as billing, bug, or feature. Return only a JSON object with label and a short evidence quote. Ticket: I was charged twice for the same invoice. Can you refund one charge? Do not take any action on the account. """ for model in ( " claude-haiku-5-5 " , " claude-opus-5-5 " ): body = { " model " : model , " max_tokens " : 4096 , " output_config " : { " effort " : " low " }, " messages " : [{ " role " : " user " , " content " : prompt }], } request = urllib . request . Request ( " https://api.beatapi.io/v1/messages " , data = json . dumps ( body ). encode (), headers = { " Authorization " : " Bearer " + key , " Content-Type " : " application/json " , " anthropic-version " : " 2023-06-01 " , }, method = " POST " , ) start = time . monotonic () with urllib . request . urlopen ( request , timeout = 90 ) as response : result = json . load ( response ) text = "" . join ( block . get ( " text " , "" ) for block in result . get ( " content " , []) if block . get ( " type " ) == " text " ) print ( json . dumps ({ " model " : model , " elapsed_seconds " : round ( time . monotonic () - start , 3 ), " stop_reason " : result . get ( " stop_reason " ), " usage " : result . get ( " usage " ), " answer " : text , }, ensure_ascii = False )) This makes two billable requests when run. It records total non-streaming request duration, not time to first token . It does not guarantee valid JSON, implement retries, or form a benchmark. Validate the returned object and inspect stop reasons before treating the result as accepted. Avoid logging sensitive tickets or credentials in production. Start a new, source-based request when escalating between models. Cross-model or cross-account replay of opaque thinking blocks is not a safe generic routing strategy; preserve those blocks only as required by the original conversation's API rules. How to run an evaluation that answers your routing question Use a held-out set with ordinary cases and deliberately hard cases. A pilot of 50–100 labeled examples can reveal integration and rubric problems; it is too small to establish a low production error rate. Measurement What to retain Why it changes the decision Accepted-result rate Expected result, rubric, observed output Distinguishes useful answers from plausible prose False acceptance Gate decision and independent correctness label Finds errors that escalation never sees Cost per accepted result Every attempt's usage and applicable rates Includes failures and rework End-to-end latency Start, finish, tool and review time Captures the actual user wait Escalation fraction Route and rejection reason Tests the routing cost formula Cache behavior Writes, reads, misses, total prompt length Separates warm-cache estimates from real traffic Keep prompts, tools, input snapshots, output limits, and grader rules fixed. Alternate model order to reduce time-of-day bias. Run a matched-effort comparison first, then allow each model a tuned configuration: those answer different questions. Record all attempts, rather than keeping only the nicest response. For a three-route experiment—Haiku only, Opus only, and Haiku with escalation—evaluate the same task IDs. The routing experiment must grade final delivered answers and record both model calls where escalation occurs. Do not compare a routed system's final success rate with a single model's first-attempt success rate without labeling that difference. For a business-facing metric, calculate: cost_per_accepted_result = (all inference + tools + validators + review cost) / accepted_results If no result is accepted, report the metric as undefined rather than zero. Rejected tasks still contributed cost. Report the acceptance rate beside this metric so a route cannot appear efficient simply by dropping hard cases. Questions worth answering before switching Is Haiku 5.5 a replacement for Opus 5.5? It is a candidate replacement for particular task classes once they meet your acceptance bar. A file scout and a migration planner can belong in the same application while requiring different models. Is Haiku always 40 times cheaper? No. That ratio applies to equal uncached input/output token counts at standard rates with prompts up to 100K. Long-prompt rates, caching mixes, generated reasoning, and retries change the comparison. Does caching keep Haiku under the 100K boundary? No. Cached input still contributes to prompt length. Count the complete prompt before choosing a tier. Does a 1M context window mean I should send the whole repository? No. It is capacity, not a retrieval strategy. Test whether a relevant slice preserves acceptance while reducing cost and latency. Keep source references available for follow-up. Can I compare API bills with a Claude subscription? Not directly. This article models per-token API rates. Subscription allowances, usage accounting, and user workflows are a separate comparison. Is Haiku faster in my application? Anthropic positions it as its fastest standard-speed model. That does not predict latency on your route or workload. Measure end-to-end completion and time to first token separately, including tools and escalation. What should I deploy first? Choose one bounded, labeled workload. Check the model ID and request shape, capture actual usage, validate results, and compare all three routes. Expand only after the candidate route meets your correctness and latency requirements. The practical decision in Claude Haiku 5.5 vs Opus 5.5 is where to spend reasoning. Give narrow, checkable work to the inexpensive candidate; spend more on unresolved complexity; and make the acceptance check strong enough to know the difference.
Christian HeilmannTop Front-end Bloggersread at source
Returning to Bengaluru in November, see you there.
After rocking San Jose with our World Congress, WeAreDevelopers’ next port of call is Bengaluru in India. See you there on the 25th of November for two days of excellent talks and workshops !
Hacker News Front PageThe Giantsread at source
Our $445M Series D
Hacker News Front PageThe Giantsread at source
Deno Is Joining Cloudflare
Hacker News: Show HNLibraries, Frameworks, etc.read at source
Show HN: UserAlertX – let your AI agent text you (only you)
Comments
Hacker News Front PageThe Giantsread at source
Study: Exercise increases cancer survival rates
HeyDesignerDeveloper/Designer Newsread at source
Tesler’s Law: Complexity moved to a different part of the journey
Designing Michelangelus, Critique in the AI era, The web needs interactive lists.
Hacker News Front PageThe Giantsread at source
Meadows – a small language for stock-and-flow diagrams that run
Hacker News Front PageThe Giantsread at source
Iranian campaign planted fake articles in real U.S. publications using ChatGPT
Hacker News: Show HNLibraries, Frameworks, etc.read at source
Show HN: A Piet interpreter compiled to a Piet image, running a Piet program
Comments
Hacker News: Show HNLibraries, Frameworks, etc.read at source
Show HN: All Things Banana – banana news, prices, and a daily Wikipedia game
Comments
Hacker News: Show HNLibraries, Frameworks, etc.read at source
Show HN: Spectra – A drop-in optimizer that cuts LLM training VRAM by 50%
Comments
Flavio CopesMore Front-end Bloggersread at source
How to install your own apps on your iPhone with Xcode
Put the apps you build on your iPhone without the App Store: turn on Developer Mode, install from Xcode, update over Wi-Fi from the terminal, and renew them.
Hacker News: Show HNLibraries, Frameworks, etc.read at source
Show HN: Long term Memory and 50M token window for LLM
Comments
SitePointThe Giantsread at source
Best SOC 2 Software in 2026: 10 Platforms Ranked by Audit Readiness
The best SOC 2 software in 2026, ranked by how fast each platform gets you from kickoff to a signed audit report - not by feature-list length. Continue reading Best SOC 2 Software in 2026: 10 Platforms Ranked by Audit Readiness on SitePoint .
PiccalilliMore Front-end Bloggersread at source
The Index: Issue #201
2027 web platform feature ranking This is the second year we're doing this, and last year's 1900+ rankings not only helped us push the right proposals in the Interop process, it was also used to prioritize web platform feature development in Firefox. Certainly beats thumbs up in GitHub issues! Make your voice heard. This hue shall pass A fabulous colour contrast tool from the great folks over at Clearleft. Dark forestry Not only an outstanding and important piece of writing, but the design and effects are lovely. Destroy any website This is very fun! Destroy your least favourite websites (like the pre-loaded example) for a vibe boost. css.earth Explore earth, other planets and even other galaxies, all rendered with HTML and CSS, via PolyCSS . The box model and box sizing Here's one from the Piccalilli archives that you might have missed to wrap up this issue. P.S. this is a good website from personalsit.es . Sponsor message Has your organisation gone too deep with frameworks or [gasp] LLMs and feel like you're stuck? We're helping clients get away from that noise and instead, focusing their efforts and budgets on actually doing stuff that will help them reach their goals. We do all of that while running this publisher, Piccalilli too! Our availability opens up again in late 2026 into 2027, so check out what we're about. Maybe we can be the key that unlocks your success in 2027 onwards. Check out our work
W3C BlogBrowsers, engines, etc.5 min read
What’s new in the W3C Website Design System
A significant update to the W3C Website Design System was published at the end of September 2026. Here’s what changed and why.
Robin WieruchMore Front-end Bloggersread at source
Toward a Self-Improving Agentic Code Review Loop
An agentic code review loop where each developer fires their own AI review skill at pull requests, the author's agent answers, and a scorecard rates each skill.
SitePointThe Giantsread at source
How to Build a Fast, Responsive Artwork Gallery for the Web
null Continue reading How to Build a Fast, Responsive Artwork Gallery for the Web on SitePoint .
AbduzeedoMulti Author Blogsread at source
Montreal Architecture Brand Identity by Mariane Farias
Mariane Farias crafts a minimalist brand identity for Montreal Architecture, translating circular window facades into tactile yellow stationery. Montreal Arquitetura is an architecture studio based in...
Frontend Masters BlogMore Front-end Bloggersread at source
How an accessibility designer adds keyboard shortcuts to a web app
Eric Bailey explains the thinking and research that go into adding just a few custom keyboard shortcuts to a web app. There’s an absolute ton to consider and test, and plenty of challenges can come from your own team.
Google for DevelopersYouTube ChannelsVideoWatch on YouTube
Building with Gemma 4
AbduzeedoMulti Author Blogsread at source
Inside the Craft of Reserva: Typography Design by Plau
Plau engineered Reserva Script as a pointed-pen typography design system for Reserva, pairing Copperplate roots with sliced terminal cuts. Reserva established its market presence through a visual voic...
SitePointThe Giantsread at source
Secure Cookies with SameSite: Lax, Strict, and Cross-Site Rules
A technical reference on configuring SameSite cookie attributes (Strict, Lax, None) alongside Secure and HttpOnly flags, with implementation patterns for Express, Next.js, and Django. Continue reading Secure Cookies with SameSite: Lax, Strict, and Cross-Site Rules on SitePoint .
SitePointThe Giantsread at source
Cloudflare Web Search API for AI Agents: Build a Grounded Worker
Build a Cloudflare Worker that grounds AI agents with live web results using the Web Search API (beta), and evaluate the trade-offs between AI Gateway, direct providers, and aggregators. Continue reading Cloudflare Web Search API for AI Agents: Build a Grounded Worker on SitePoint .
SitePointThe Giantsread at source
TypeScript Generics Examples: Fetch Wrappers, Forms, and UI Components
Build typed fetch wrappers, form validators, and polymorphic React components using TypeScript generics, with runtime validation boundaries and a decision table comparing generics, unions, and schema inference. Continue reading TypeScript Generics Examples: Fetch Wrappers, Forms, and UI Components on SitePoint .
SitePointThe Giantsread at source
React Dialog Exit Animation: Fixing Broken Transitions on Unmount
Diagnose why conditional rendering breaks exit animations in React dialogs, then evaluate delayed unmount states, Motion's AnimatePresence, Radix UI, native dialogs, and View Transitions. Continue reading React Dialog Exit Animation: Fixing Broken Transitions on Unmount on SitePoint .
SitePointThe Giantsread at source
Playwright 1.64 WebMCP Testing: How to Inspect and Call Page Tools
Test pages exposing WebMCP tools using the page.webmcp API introduced in Playwright 1.64. Learn how to launch Chromium with the required flag, discover page-registered tools, execute calls, and assert results. Continue reading Playwright 1.64 WebMCP Testing: How to Inspect and Call Page Tools on SitePoint .
SitePointThe Giantsread at source
Claude Haiku 5.5 Pricing: How to Calculate Production Costs and Migrate
Model production workload costs under Claude Haiku 5.5's two-tier rates, account for the 100k-token pricing threshold, and plan a phased migration from Haiku 4.5. Continue reading Claude Haiku 5.5 Pricing: How to Calculate Production Costs and Migrate on SitePoint .
Bram.usMore Front-end Bloggers4 min read
<calc-input>, a custom input element that accepts mathematical formulas
Continuing my streak of small form-related Custom Elements (like ) , I built . It’s a custom input element that accepts mathematical formulas — such as 2 + 3 or (2 + 3) * 4 — and automatically toggles between showing the formula on focus and the calculated result on blur.
Bram.usMore Front-end Bloggers2 min read
caniname: CLI tool to check project name availability across Netlify and NPM (and other sources)
Whenever I build a new tool or custom element (like or ), I always end up doing the same manual dance: checking if the package name is still free on NPM and if the matching .netlify.app subdomain is still unclaimed on Netlify. To automate that check, I built caniname .
AbduzeedoMulti Author Blogsread at source
How Brand Brothers Crafted Brand Identity for Glénat
Brand Brothers engineered a modular brand identity for French publisher Glénat, creating a geometric visual language from the historic mark. Founded in Grenoble in 1969 by comic enthusiast Jacques Glé...
Jim Nielsen’s BlogMore Front-end Bloggersread at source
“Getting off the Modernization Treadmill”
My notes from this talk by Alexander Petros at Big Sky DevCon 2026 . Alex starts by noting how “modernize” used to mean something along the lines of “update this thing that was made before I was born”. But now “modernize” means something more like “update this thing from 5-10 years ago” (hence the framing of the talk, the “modernization treadmill”). Using a real-world example of an incredibly slow website that was required to access state-sponsored programs for welfare, Alex points out the disparity in conditions between those of us who make software and those who have to use them: The people who develop these websites are usually doing so on high-powered internet connections and high-powered devices, but they're not using them in the conditions that the people who most need those benefits are going to be. Then he shows a Reddit thread where somebody essentially posted, “I’m having problems with this website. I’ve been waiting for months for my application to go through. Any suggestions?” And one Reddit user responded, “The best thing you can do is go into the physical office, get a case worker, and your problems will be solved within the hour.” The irony. I guess we've come full circle now. It used to be: “Don’t talk to anybody. It’s faster and more convenient to use the website!” But now it’s: “Don’t use the website. It’s faster and more convenient to talk to somebody!” Have we failed at making websites? And is our failure, at least in part, rooted in the fact that we don’t leverage the basic tools for making websites: HTML, CSS, and (in a distant third) JavaScript? Alex goes on to argue that the technologies of the web have an ideological bent and, if used as designed, can solve so many of the performance, accessibility, and usability issues that plague so many websites. The grain of the web’s technologies are rooted in these values: User-friendly Backwards- and forwards-compatibility Long-term viability Universal accessibility Which means if you use them as intended, they are optimized to deliver outcomes rooted in those same values. So if you like those values and you want those outcomes, use the platform. Take HTML, for example. Here’s Alex: HTML does [performance improvements] for you for free. If you've coded your website in a proper, semantical, structure HTML style, it will just get better over time at zero cost to the people who built that website HTML is your friend. HTML won’t give you up or let you down . HTML will make it difficult for you to make a bad website. Write it in to the requirements of the project you’re doing that it work without JavaScript. Not necessarily that it doesn’t have any JavaScript, but just that the core functionality of the website can happen without JavaScript. If you do this […] you will find that it’s very hard to deliver a bad web page because the structure that HTML requires is one that fundamentally is good for the user, performant, and cost effective. Technologies are imbued with culture, which influences what you do and how you do it. If you can align your ideological beliefs with your technological choices, you might end up with an outcome that aligns with your values — who would’ve thought, eh? A lot of modern software developers come from websites like Facebook, they come from big tech companies [who] fundamentally have a different set of priorities. Their job is to keep you on the website as long as possible so that you consume more ads. But that’s the opposite set of requirements and priorities that the government needs to be doing, which is to build something that is clean, quick, efficient, and gets you in and out as fast as possible. So [I tell people] that the technology they use comes from [an] ideological place. But there are different ideological places that produce different technological results and if we start from those through lines then we can produce services that help people who need them. The ideological principles of the web are well established : users over everything else. If you use HTML as much as possible, you’ll make something that’s as user-friendly as possible on the web. Reply via: Email · Mastodon · Bluesky
Next.js BlogLibraries, Frameworks, etc.read at source
Upcoming Next.js Security Update for Upstream Vulnerabilities
Next.js plans to publish an out-of-band security update next Wednesday, October 14, 2026, addressing two Critical and one High severity vulnerabilities in upstream dependencies.
Google for DevelopersYouTube ChannelsVideoWatch on YouTube
📷 Searching for images from text with EmbeddingGemma 2
Web Tools WeeklyMulti Author Blogsread at source
Web Tools Weekly Issue #690
Media Tools, JS Plugins, Git/CLI Tools
AdactioTop Front-end Bloggersread at source
The people holding up the internet | Data Drop
Bringing receipts for xkcd.com/2347 . adactio.com/links/22794
AdactioTop Front-end Bloggersread at source
A Place Among the Fiddlers - Longreads
An in-depth report from Galax, the Old Time equivalent of the fleadh. adactio.com/links/22793
AbduzeedoMulti Author Blogsread at source
Off The Post #01 — Editorial Design by Laís Zanocco
Laís Zanocco and Estúdio Drama craft dynamic editorial design for Off The Post, pairing collage storytelling with bold sports typography. Most sports magazines look the same. They rely on cold photo g...
Kilian ValkhofMore Front-end Bloggers2 min read
I’d like to have a universal pseudo selector in CSS
Last week, I wasted about an hour of my time figuring out why a certain element had different dimensions than the same element elsewhere in the DOM. It had the same CSS and the same defined width and height, but it still rendered differently. Yes, it was box-sizing. Of course it was box-sizing. I immediately […] The post I’d like to have a universal pseudo selector in CSS first appeared on Kilian Valkhof .
freeCodeCamp.orgYouTube ChannelsVideoWatch on YouTube
TanStack Router Course – Loaders, Auth & Type-Safe Routes
HeyDesignerDeveloper/Designer Newsread at source
Context is now the product (video)
Fluid functionalism, Design owns coherence, Imperfection.
W3C NewsBrowsers, engines, etc.read at source
First Draft Notes: Verifiable Credential Threat Models
The Verifiable Credentials Working Group has published a series of W3C Working Group Note Drafts for threat models.
Flavio CopesMore Front-end Bloggersread at source
How to use Jev with Claude Code, Codex and Cursor
How to use Jev with Claude Code, Codex and Cursor: install TypeSafe's agent skill, give the agent the docs, and review the Jev code it writes.
Flavio CopesMore Front-end Bloggersread at source
The complete guide to LiveKit Agents
Build a voice AI agent with LiveKit Agents in Node.js and TypeScript, then test it, simulate calls, deploy it to LiveKit Cloud and watch real sessions.
The Mozilla BlogBrowsers, engines, etc.1 min read
Firefox partners with Anschutz Entertainment Group at the Uber Arena and Uber Eats Music Hall in Berlin
Firefox is teaming up with AEG and taking to the main stage at Berlin’s biggest and best indoor arenas. Starting this October, as the indoor music scene heats up and the basketball and ice hockey seasons are in full swing, fans flocking to the Uber Arena, Uber Eats Music Hall, and Uber Platz will see […] The post Firefox partners with Anschutz Entertainment Group at the Uber Arena and Uber Eats Music Hall in Berlin appeared first on The Mozilla Blog .
AbduzeedoMulti Author Blogsread at source
Timeless Unveils Interactive Web Design for Timeless Type
Timeless crafts an interactive web design for Timeless Type, pairing tactile 3D controls with a four-family variable font system. Digital type specimens and foundry landing pages often fall into predi...
Ana RodriguesMore Front-end Bloggers1 min read
My Interop feature ranking
The Firefox team just launched a website that allows users to rank the Interop 2027 proposals that they care about. And I care about things - for selfish reasons and for work. And that's a challenge. I want to rank things that only benefit me and my blog but I also want to rank…
Bootstrap BlogLibraries, Frameworks, etc.read at source
Bootstrap 6 Alpha
DebugBear BlogCompany/Startup Blogs28 min read
Astro Tutorial: Make a Static Site with Interactive Components
This Astro tutorial shows how to build an interactive blog using file-based routing, build-time and live content collections, and the islands architecture.
Swizec TellerMore Front-end Bloggers1 min read
When code is cheap, judgement becomes the job
Software engineering has changed in the last 2 years. Here's how that looks without the fluff. Real company, big stakes, finite token budget
Twilio BlogMulti Author Blogsread at source
Building a Multi-Channel AI Agent with Twilio Conversations and eve
Learn how to build a multi-channel AI agent with Vercel's eve and Twilio Conversations. Unify SMS and WhatsApp into one thread, give your agent long-term customer memory with Twilio Memory Store, and swap the LLM with one line of TypeScript.
Twilio BlogMulti Author Blogsread at source
[Webinar] Master automated journeys
Move beyond unwanted communications. Drive immediate ROI and hyper-relevant marketing.
Twilio BlogMulti Author Blogsread at source
Unified Customer Communications: Twilio Voice, SMS, and WhatsApp meet Salesforce
Unify Twilio Voice, SMS, and WhatsApp directly in Salesforce with the new Partner Contact Center Integration. Eliminate the tab-swapping now.
Google for DevelopersYouTube ChannelsVideoWatch on YouTube
Real-time golf swing coaching with Gemini 3.8 Live
Jens Oliver MeiertTop Front-end Bloggersread at source
On the Decision Underlying Decorative Images
What would change the judgment people have to make about images?