Skip to content
Frontend Feeds
  • Today
  • Archive
  • Sources
  • Categories
    • The Giants6 Sources
    • Multi Author Blogs18 Sources
    • Top Front-end Bloggers25 Sources
    • More Front-end Bloggers77 Sources
    • Browsers, engines, etc.11 Sources
    • Libraries, Frameworks, etc.12 Sources
    • Company/Startup Blogs12 Sources
    • Developer/Designer News4 Sources
    • YouTube Channels13 Sources
    • Podcasts7 Sources

The Giants

DEV Community

5 itemsVisit site ↗

Sunday, 4 October 2026

  • DEV CommunityThe Giantsread at source

    BottleBuddy

    This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built BottleBuddy , a small AI-powered product discovery app for a friend. The idea came from a simple problem: browsing a large list of products and manually applying filters can be annoying when you already know roughly what you want. Instead, BottleBuddy lets my friend describe what they're looking for in natural language. For example: "I'm looking for beer under Rs. 5,000 (LKR) with a lower ABV." BottleBuddy interprets that request and converts it into structured filters such as budget, category, brand, and ABV range. It then searches the product data stored in MongoDB and displays matching products. I also included regular filters so the application isn't dependent entirely on AI. Demo BottleBuddy currently runs locally because the AI model is running directly on my machine through Ollama. Here's a video showing the complete flow: https://youtube.com/shorts/uuS9bb2CIQA?feature=share In the demo, I show: entering a natural-language request Gemma interpreting the request BottleBuddy finding matching products product information coming from MongoDB manual filtering/search functionality Code GitHub repository : https://github.com/ImalKesara/BottleBuddy How I Built It I built BottleBuddy using SvelteKit and TypeScript for the application. The AI part uses Gemma 3 , Google's open-weight model, running locally through Ollama . When someone enters a request such as: "Beer under Rs. 5,000 lkr, preferably lower strength" BottleBuddy sends the text from the SvelteKit server to the locally running Gemma model. Gemma's responsibility is not to invent or recommend products. Instead, it interprets the natural-language request and converts it into structured search filters. For example: { "maxBudget": 5000, "category": "beer", "brand": null, "minAbv": null, "maxAbv": 5 } BottleBuddy then uses those filters to search the actual product data stored in MongoDB Atlas . Why Does Open Innovation Matter? Open innovation made the core idea behind BottleBuddy possible. Because Gemma is an open-weight model, I was able to run the model locally on my own machine using Ollama and integrate it directly into my application. That gave me much more control over how the AI part of the application works. The application is not tied to a closed AI API for its core natural-language search feature, and the model can be run on infrastructure I control. It also made experimenting much easier. I could test prompts, structured outputs, and different ways of extracting search filters while keeping inference local. For BottleBuddy specifically, this showed me that useful AI features don't always require sending requests to a large proprietary hosted service. An open-weight model running locally can be enough to build the core experience.

  • DEV CommunityThe Giantsread at source

    Model Evaluation in Machine Learning: How Do We Know a Model Is Good?

    Model Evaluation in Machine Learning: How Do We Know a Model Is Good? Building a machine learning model is only one part of a machine learning project. After training a model, we need to answer an important question: How well does it actually perform on new data? This is where model evaluation becomes important. What Is Model Evaluation? Model evaluation is the process of measuring how well a machine learning model makes predictions. It helps us understand whether a model is performing correctly and whether it can generalize to data it has never seen before. A model that performs extremely well on training data may still perform poorly on new data. This problem is commonly known as overfitting. Proper evaluation helps identify such problems before using the model in a real-world application. Classification Model Evaluation For classification problems, several metrics can be used. Accuracy represents the percentage of predictions that the model classified correctly. Although it is easy to understand, accuracy may not be reliable when the dataset is imbalanced. Precision tells us how many of the observations predicted as positive were actually positive. Recall measures how many of the actual positive cases were correctly identified by the model. F1-score combines precision and recall into a single metric. It is particularly useful when both false positives and false negatives are important. Another useful tool is the confusion matrix, which shows true positives, true negatives, false positives, and false negatives. Regression Model Evaluation For regression problems, where the model predicts continuous values, different metrics are commonly used. Mean Absolute Error (MAE) measures the average absolute difference between predicted and actual values. Mean Squared Error (MSE) calculates the average squared difference between predictions and actual values. Root Mean Squared Error (RMSE) is the square root of MSE and is useful because its value is expressed in the same units as the target variable. The R² score indicates how well the model explains the variation in the target variable. Why Cross-Validation Matters Another important evaluation technique is cross-validation. Instead of relying on a single train-test split, the dataset is divided into multiple parts. The model is trained and evaluated several times using different portions of the dataset. This gives us a more reliable estimate of model performance and helps reduce the risk of making conclusions from one particular train-test split. Final Thoughts Model evaluation is essential for building reliable machine learning systems. Choosing the right evaluation metric depends on the type of problem and what we want the model to achieve. Accuracy, precision, recall, F1-score, confusion matrices, MAE, MSE, RMSE, R², and cross-validation are some of the most useful tools for evaluating machine learning models. For a more detailed explanation of Model Evaluation, you can read my related article

  • DEV CommunityThe Giantsread at source

    One field in the request made our agent 3x cheaper and 8x faster

    TL;DR. Reasoning models decide by themselves how long to think if you don't tell them. Our agent didn't — and on hard tasks the model sometimes thought for 14 minutes and 33K tokens in a single step. One field in the request body ( "reasoning": {"effort": "low"} ) made a task 3x cheaper and 7–8x faster, and it solved more tasks, not fewer: 12 of 12 instead of 6 of 7. The best setting in the end was neither "always low" nor "always high" but "low, and high right after a failing test". Below: how we measured it, the tables, the code, and what it doesn't fix. About this post. The project is Altair , an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the experiments and the charts were prepared by that same AI assistant; I reviewed it and stand behind it. Every number comes from our runs. This is the third post about the project: the first is about snapshots and running tests before "done", the second about cutting a browser agent's tokens by 58%. How it started We ran the agent on a cheap reasoning model ( glm-5.3-flash ) and read the provider's logs. On one task the very first step took 14 minutes. The provider had time to bill it: 32,978 reasoning tokens in one step , and no answer. Our loop guard cut the stream at 160,000 characters. The cause was mundane. Reasoning models have a "how much to think" knob — reasoning_effort at OpenAI, reasoning.effort at OpenRouter, thinking in Z.ai-style APIs. If you don't send it, the provider or the model decides; OpenRouter's docs say as much. Our agent sent nothing. How we measured Easy tasks (fix a bug, answer a question about code) were solved 100% of the time under any setting, so they show no difference. We wrote six hard tasks with hidden tests — the agent never sees them, they run after it says "done": an LRU cache with a time-to-live (11 hidden edge-case tests); fixing a config parser from seven user bug reports; renaming a function across a multi-file package and adding a parameter; the 95th percentile of latencies from logs, with exclusions; parsing "1.5M", "200K", "10 тыс"; business days between dates with holidays, fast over a hundred years. Each task was first solved with reference code to make sure the hidden tests were fair. The agent is the real Altair, run through its CLI ( altair -p ). A proxy between agent and provider logged every request: size, cached tokens, reasoning tokens, time. Cost is what OpenRouter billed. Results Cost of one hard task Setting Solved Cost per task Time per task No level (before) 6 of 7 12.6 m$ 4.1 min effort = low 12 of 12 4.2 m$ 0.55 min effort = medium 10 of 12 4.8 m$ 0.64 min effort = high 12 of 12 6.4 m$ 1.5 min m$ is a thousandth of a dollar. 12.6 m$ without a level, 4.2 m$ with low : three times less. Time went from 4.1 to 0.55 minutes, 7.5x. Why "6 of 7" and not "of 12": runs without a level kept looping for 10–14 minutes, and we stopped that series early so as not to burn money. One of the six tasks (parsing "1.5M") sent the model into endless reasoning in almost every series without a level. With any explicit level — never. Where the money goes: Reasoning tokens per task 9,743 reasoning tokens per task on average without a level, 235 with low — 41x fewer. Output tokens cost 3.3x more than input for this model, so on hard tasks reasoning is the main bill. And on long tasks? Short tasks are 5–9 steps. We also tried 15–20-step ones: a package with six bugs in different modules, auditing twelve values across sixty files of our own code, and "read eight files in full, then answer questions about them". Long tasks 9 of 9 solved in every setting, but without a level a task costs 28 m$ and takes 5.2 minutes; with an explicit level, 11–13 m$ and a little over a minute. The twist: think hard only when something broke low is cheap, high is safer. We wanted both, and tried two per-step modes: "think about the plan" — high on the first step, then low ; "think after a failure" — low , but high on the step right after tests failed or a tool returned an error. When to think harder Setting (18 runs each) Solved Cost Time always low 16 of 18 3.78 m$ 0.96 min always high 17 of 18 5.72 m$ 1.92 min high on the first step 15 of 18 4.37 m$ 1.93 min low , high after a failure 17 of 18 3.65 m$ 0.84 min Careful planning didn't help — it solved the fewest. "Think after a failure" solved as many as always- high , at the cost and speed of low . It makes sense: most of an agent's work is routine (read a file, write a file, run the tests), and thinking pays off where a test just showed the first idea was wrong. One task in 18 is close to noise, honestly, but this mode is no worse than low on all three measures. It's now our default. How it looks in code The catch is that every provider has its own field: def reasoning_extra ( level , base_url , model , override = " auto " ): """ Request fields for a reasoning level ({} for " default " / unknown levels). """ if level not in LEVELS : return {} kind = dialect ( base_url , model , override ) if kind == " openrouter " : # OpenRouter takes minimal..high; some endpoints refuse "none"/enabled=false, # so the lowest we send is "low" for "minimal" there. return { " reasoning " : { " effort " : " low " if level == " minimal " else level }} if kind == " zai " : # Z.ai-style APIs have an on/off switch: low and minimal turn thinking off. return { " thinking " : { " type " : " disabled " if level in ( " minimal " , " low " ) else " enabled " }} if kind == " openai " : return { " reasoning_effort " : level } return {} What we stepped on along the way: For this model OpenRouter refuses reasoning: {"enabled": false} and reasoning_effort: "none" with a 400 "Reasoning is mandatory for this endpoint". Only effort: "low" works. A Z.ai-style gateway, the other way round, takes thinking: {"type": "disabled"} and didn't answer reasoning.effort at all. A provider that doesn't know the field may answer 400. We catch that, retry without the field and remember not to send it again. The agent loop picks the level for each step: FAILURE_RE = re . compile ( r " \b\d+ failed\b|\bFAILED\b|Traceback \(most recent call last\)|AssertionError| " r " \bexit code [1-9]\d*\b|\bSyntaxError\b|[A-Za-z]Error: " ) def _reasoning_level ( self ) -> str | None : mode = self . settings . llm_reasoning if mode == " default " : return None if mode == " adaptive " : return " high " if getattr ( self , " _round_failed " , False ) else " low " return mode After each round of tools, _round_failed is true if a tool returned an error or its output has "2 failed", a traceback or a non-zero exit code. Service calls thought more than the main ones Two more places turned up in the logs. The chat title. The agent asks the model for a title from the first message. About 95% of that answer was reasoning: ~190 tokens of "thoughts" for a five-word title. With reasoning off the answer is 12 tokens and comes 1.6–3x faster; all six sample titles were still fine. The history summary. When the context grows, the agent asks the model to condense the start of the conversation. The answer limit was 600 tokens. The model spent them on reasoning and the summary came back empty in 3 cases out of 4 ( finish_reason: length ). An empty summary isn't a fold, it's a loss: the agent drops the start of the conversation instead of condensing it. With a 2,000 limit there were no empty summaries. Service calls now ask for the least reasoning, and the summary limit is 2,500. What this doesn't fix One model. Everything was measured on glm-5.3-flash . Other models scale their levels differently, and "low" may mean something else. The principle — don't leave the level to the provider — carries over; the exact numbers don't. Small samples. 12–18 runs per setting. A difference of one or two solved tasks is noise. A difference of 3x in cost and time is not. Not every provider has levels. A Z.ai-style gateway only knows on/off. There "high after a failure" means "think fully", which is expensive: on our tasks "always off" on that gateway came out half the price of adaptive (5.9 vs 10.8 m$, 12 of 12 both). If cost matters most, pick "low" there. The tasks are code. For writing, search or analytics the best level may differ. Check it yourself The core of the experiment is one agent run with a set level: # proxy.py adds the level to every agent request to the provider body = { ** { " reasoning " : { " effort " : " low " }}, ** request_body , " model " : provider [ " model " ]} # stand.py: the agent solves the task, then the hidden tests run subprocess . run ([ sys . executable , " altair_cli.py " , " -p " , " --output-format " , " json " , " --mode " , " bypass " , " --cwd " , workspace , task . prompt ], env = env , timeout = 900 ) passed , detail = task . check ( workspace , answer ) The whole stand (the logging proxy, tasks with hidden tests and reference solutions, the summary script) and the raw results: research/agent-lab-2026-10 ( python lab/facts.py recomputes every number in this post from the saved results, no keys needed). Altair's code: https://github.com/Qweezyy/AltairAgent — the level is in Settings → Agent → Reasoning level (default "Adaptive"), LLM_REASONING in .env , module pc/core/llm/reliability.py . How do you set the reasoning level in your agents — fixed, per task type, or as the work goes? Have you seen a model think far more than it needs when no level is set?

  • DEV CommunityThe Giantsread at source

    I Built a Marketplace for Developers to Buy Ready-to-Use Source Code

    I Built a Marketplace for Developers to Buy Ready-to-Use Source Code I've been building full-stack projects for the past few years, and I kept running into the same problem: Sometimes you want to build a SaaS, dashboard, mobile app, API, or developer tool — but you don't necessarily want to start every project from an empty folder. You might want a solid starting point, a codebase to customize, or simply something real to study. That idea led me to build Projectry . 🌐 https://projectry.dev Projectry is a marketplace for developers looking for ready-to-use source code and complete projects that they can learn from, customize, and build on. What’s currently available I started with six projects covering different use cases and technology stacks: 📚 StudyStack An education platform focused on flashcards, spaced repetition, and learning progress. Stack: React, TypeScript, Tailwind CSS, MongoDB 📊 DevBoard A DevOps dashboard for monitoring deployments, CI runs, services, and incident timelines. Stack: Python, FastAPI, MongoDB, React 📋 TaskForge A project management application with Kanban boards, collaboration, activity feeds, and role-based workspaces. Stack: Next.js, TypeScript, Node.js, PostgreSQL 📝 PaperTrail An audit logging and compliance tracking system with a review dashboard and tamper-evident storage. Stack: Java, Spring Boot, PostgreSQL, Docker 📦 ShelfLife An Android inventory and expiry tracking application with barcode scanning and offline capabilities. Stack: Kotlin, Android, FastAPI, PostgreSQL 🔐 CodeVault A developer-focused encrypted snippet and secret vault with versioning, sharing, and an asynchronous audit pipeline. Stack: Java, Spring Boot, PostgreSQL, Docker, RabbitMQ Why build another source-code marketplace? I didn't want Projectry to be just another collection of generic templates. The goal is to build a broader collection of real-world applications across different areas: Backend systems Full-stack applications Developer tools DevOps Education Mobile applications Data applications Frontend projects The catalog also covers technologies such as React, Next.js, TypeScript, Java, Spring Boot, Python, FastAPI, PostgreSQL, MongoDB, Docker, RabbitMQ, Redis, Android, and more. What's next? The six projects are just the beginning. I'm working on expanding the catalog while keeping the projects varied in both technology and use case. Some of the areas I'm exploring include: SaaS applications E-commerce Developer infrastructure Data dashboards CI/CD tooling APIs and backend systems Mobile applications Business applications Authentication and authorization Real-time applications I'm also working on better documentation, demos, and developer resources around the projects. The idea is to provide more than just a ZIP file and leave you to figure out the rest. For developers If you're looking for a starting point for a project, want to study how a real application is structured, or need a codebase that you can customize instead of building everything from scratch, check out Projectry: 👉 https://projectry.dev I'm also building the Projectry GitHub organization, where I'll be publishing selected resources and open-source projects: 👉 https://github.com/Projectry-Org I'm still early in the journey, so I'd genuinely like to hear from other developers: What kind of source code would you actually want to buy? A complete SaaS? A backend/API? A mobile app? A dashboard? A developer tool? Something else? I'd love to hear what you'd find useful.

  • DEV CommunityThe Giantsread at source

    126+ Private, In-Browser Web & PDF Tools Every Developer Needs

    Most online converters and PDF tools upload your sensitive documents, API keys, and contracts to remote cloud servers. For private data, that is a security risk. I built DevTol — a free suite of 126+ developer and productivity tools that execute 100% inside your browser using Web Workers and the Web Crypto API. What is inside DevTol? Full Client-Side PDF Suite: Edit PDFs (with full Arabic RTL support), Compress PDFs locally (down to 100kb/200kb), Merge, and check Digital Signatures. Developer Utilities: JSON Diff & Formatter, JWT Decoder, Regex Sandbox, SQL Prettifier, cURL to Code Converter (7 languages). AI Token Calculator: Calculate prompt tokens and API pricing across GPT-4o, Claude 3.5, Gemini, and DeepSeek. Zero Tracking: 100% free, no account needed, and no files ever leave your machine. Try it out here: https://devtol.online What tools would you like added next? Feedback is welcome!

Articles belong to their publishers. This site only collects what their feeds provide.

Updated 4 October 2026 at 11:35