New Oct 1, 2026

Under the Hood: Prompt tuning Shake to Summarize for recipes

Browsers, engines, etc. All from The Mozilla Blog View Under the Hood: Prompt tuning Shake to Summarize for recipes on blog.mozilla.org

Following the rollout of Firefox’s Shake to Summarize feature to Android devices this May, we wanted to take a closer look at the modeling work behind the feature. In a previous blog post, we discussed our model selection process. Now that Shake to Summarize has been available on both iOS and Android for a few months, we wanted to take it a step further and outline our approach to prompt development. 

This is the story of a Shake to Summarize use case that required a little extra prompt tuning: online recipes.

Framing the Problem

The first step in developing a useful prompt is clearly describing what you want the LLM to do. LLMs thrive on specificity, so the more sharply you can define your task, the better your results are likely to be. 

This is especially important when using smaller LLMs (like we are here), since these little models are not as good at reading between the lines and intuiting unstated intentions as their more powerful cousins are. 

For this application, we were looking to create summaries. On the face of it, this seems pretty straightforward. However, as we iterated on prompts, we quickly discovered that what constitutes a “good” summary depends largely on what one is summarizing. 

For example, a useful summary of a novel should provide us with a quick overview of the plot without getting into too many specific details; we wouldn’t expect the summary to contain anything from the text verbatim. 

In contrast, a summary of a recipe website should include the recipe essentially as written. If the recipe says “cook lovingly” we might be OK with shortening it to “cook,” but if it calls for 4 cups of vegetable broth, a Tbsp of oregano and a tsp of thyme, we want these details relayed to us exactly. Merely stating, “this recipe calls for some broth and some spices” would not be adequate. 

Forming the Prompt

From here, it became clear that we wanted not just one “summarize” instruction, but a whole set of instructions — one for each category of webpage that we wished to summarize. We worked with the product team to compile a list of article types that we were targeting for this feature. For each article type, we then gave a brief description of what a good summary would look like: 

Recipe – Ingredients as written, along with the key steps, time required, and any tips given by the author or commenters

News – Only the important details: what happened, when, and what would be the likely consequences to the reader

How-to – Start with the required materials, skills, tools, etc.,  along with the main steps and any specific warnings called out by the author

Review – Highlight the bottom line rating. If it is a product review, include the pros and cons, the price, and who the target audience is.

Research – Key finding and level of confidence the researchers have in their results as well as the projected real-world impact.

Opinion – The main argument, along with any key evidence cited in support

We then wrapped this list in some general instructions to ground the model in the task and came up with our initial prompt (formatted for readability) [source]:

You are a Content Summarizer. You create mobile-optimized summaries by first understanding what users actually need from each type of content. 

Process: 

   Step 1: Identify and Adapt.

   Step 2: Based on content type, prioritize: 

Recipe – Servings, Total time, Ingredients list, Key steps, Tips.

News – What happened, when, where. 

How-to – Total time, Materials, Key steps, Warnings. 

Review – Bottom line rating, price.

Opinion – Main arguments, Key evidence.

Personal Blog – Author, main points. 

Fiction – Author, summary of plot.

All other content types – Provide a brief summary of no more than 6 sentences.

   Step 3: Format for mobile using concise language and paragraphs with 3 sentences maximum.

Bold critical details (numbers, warnings, key terms).

We Foxfooded the above prompt and found that it mostly worked pretty well. The summaries the model produced were concise, accurate, and informative. 

The only exception was recipes. 

We found that, when the prompt above was used for recipes, the model tended to summarize too much — often leaving out key ingredients, or in some cases, leaving out the recipe entirely.

For example, when asked to summarize a lentil soup recipe the model returned:

The webpage provides a detailed recipe for a vegan lentil soup, which the author claims is the best they’ve ever made. The soup is made with pantry staples like lentils, diced tomatoes, vegetable broth, and spices (cumin, curry powder, thyme), along with fresh vegetables (onion, carrots, garlic) and greens (kale or collard greens). A key step is blending a portion of the soup to create a creamy texture without dairy.

The summary is accurate, so far as it goes, but it isn’t exactly useful. If our goal is to have the summary be a quick way to access the recipe itself, without having to read through the narrative preamble, this wouldn’t cut it.  

Focusing the Prompt

To address this problem, we needed to be more explicit with the instructions we gave the model when summarizing recipes: we needed to tell it which information we were interested in and which it could safely discard. 

Although we could have added this extra guidance directly to the prompt above, doing so risked causing the model to over-index on the recipe formatting and ignore the other article types. Instead, we developed a separate prompt which contained only the recipe instructions and directed recipe summarization requests to use it, rather than the general one [source].

To perform this routing, we used the structured data embedded in each webpage. While this categorization could have been handled by another model pass, metadata from the webpage itself allowed us to quickly and deterministically identify the page categories without additional inference overhead. To capture cases where this data is either missing or inaccurate, we also maintain some lightweight recipe guidance in the general prompt. 

After this change, the summary returned by the model for the above recipe became much more usable (shortened for brevity):

Servings: 6

Total Time: 50 minutes

Prep Time: 10 minutes

Cook Time: 40 minutes

🥕 Ingredients

📋 Instructions

  1. Warm the olive oil in a large Dutch oven or pot over medium heat.
  2. Add the chopped onion and carrot, cooking until the onion softens and turns translucent, about 5 minutes.
  3. …

⭐ Tips

🥗 Nutrition


With this change in place, we ran a quick test over a curated set of recipe sites and found that the new system was more than twice as likely to return a complete and accurate summary than our previous one. Success!

The system was now working as expected: summaries were useful and recipes were complete.

Reflections

From this experience we learned that the model produced the best results when it was told explicitly what we wanted it to do. When the instructions were vague or left too much up to the model, performance suffered. 

To this end, we found that framing this problem as a routing problem — where the specific kinds of articles are routed to specific prompts — worked well. Since the task of summarization is not monolithic, our summarization pipeline should not be either.

Even though our current approach has only a single category-specific prompt, we hypothesize that the system would see further gains by using dedicated prompts for other page types as well. 

More broadly, this experience reinforced an important lesson for us: improving AI systems is not solely about building larger or more capable models. Some of the biggest gains come from reducing ambiguity, narrowing the task, and designing systems that help the model succeed. Within the right harness, smaller, open source models can deliver great value.

While building more capable models continues to advance the field, our experience shows that thoughtful system design and solid engineering still matter.

The post Under the Hood: Prompt tuning Shake to Summarize for recipes appeared first on The Mozilla Blog.

Scroll to top