《费城问询报》开发了 Scrape,一款用于挖掘超本地新闻的人工智能工具
The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news

原始链接: https://www.lenfestinstitute.org/solutions-resources/philadelphia-inquirer-scrape-ai-hyperlocal-news/

《费城问询报》开发了一款名为 Scrape 的 AI 工具,用于从市政会议记录、学校日程和 Facebook 页面等分散的本地社区信息源中发掘值得报道的新闻线索。其目标是减少编辑每周约 12 小时的研究时间,同时保留编辑的判断力。 Scrape 围绕编辑室现有的工作流程构建,并以编辑手工整理的选题清单作为训练和评估基准。在三到四个月里,编辑持续标注 Scrape 每天发送的邮件摘要,识别有价值的线索、无效信息和遗漏内容。这些反馈不断优化提示词并改善结果。 团队还将编辑部的具体背景纳入系统,包括每份新闻邮件的受众、地域范围、学区、常见主题、排除项和读者兴趣。这使输出结果比仅依靠地域范围或通用的“新闻价值”标准更具相关性和可操作性。 Scrape 已成为六家、并即将成为八家本地新闻邮件的重要基础设施,共有近 20 名记者使用。主要经验包括:从真实的工作流程入手;以现有成果为基准评估 AI;让编辑持续参与审核;明确记录编辑判断标准;从一开始就为每份新闻邮件设计专门的简报;并尽早确定成本与检索质量之间可接受的取舍。

相关文章

原文
Case Study

A project of the Lenfest AI Collaborative and Fellowship Program

By Sonali Verma

September 30, 2026

The problem 

The Philadelphia Inquirer publishes hyperlocal newsletters covering specific communities across the Philadelphia area. The team needed a way to reliably surface newsworthy stories from each community from a fragmented ecosystem of sources, such as municipal meetings, school calendars, and Facebook pages. 

One of their editors was spending about 12 hours a week just finding items for newsletters, leaving little time for writing and shaping the voice of the publication. 

Gathering news for these regions was hard because it comes from “a lot of really small sources, and so it’s a lot to sift through,” said Kevin Hoffman, The Lenfest Institute’s AI Fellow who is embedded at The Inquirer.

The solution 

The Inquirer’s AI tool, Scrape, was created to save time and raise the signal-to-noise ratio of that work.

Scrape worked because it was built around The Inquirer’s workflows, with tight feedback loops from the editor, rather than as a purely technical experiment.

First, the team grounded the tool in the existing manual process. The lead editor had already compiled a curated list of sources and was producing daily tip sheets by hand. Kevin turned that into a direct AI training and evaluation loop.

Their hypothesis was that AI could be used to find newsworthy information. To test it, the team would “run AI every day over the sources that this editor cared about, produce a tip sheet, and then also compare this tip sheet to what this editor produced and understand, ‘Okay, is it actually extracting the same insights? Is it maybe finding new insights that the editor missed, or is it or is it missing in some of those areas?’” Kevin said. 

“And that is exactly what we did over the span of three or four months. It was pretty intense. Every other day, our editor would give me feedback on how Scrape was doing.”

The team would run Scrape and produce an e-mail digest that the editor would annotate, pointing out what was useful and what was not. 

This iterative prompt refinement, where bullet points were added and removed and newsworthiness was defined, gradually shifted the output from noise (such as rescheduled meetings and routine road closures) to genuinely impactful community news, Kevin said.

The second major success factor was embedding editorial context directly into the system. The team used newsletter-specific descriptions, such as the newsletter’s audience, relevant geographical regions, school districts and topics, rather than simply using geography as a guideline for hyperlocal news. In this way, Scrape’s outputs were aligned with what newsletter readers actually cared about.

“For instance, Chester County is one of our newsletters. The editor for the project composed a description for this newsletter that included information on what they were looking for – readers are interested in everything from municipal and local governments to localized state news relevant to them, especially things that will impact their lives, like restaurants and retail businesses,” Kevin said. The prompts explicitly state what not to include, such as straight business news on topics like the stock exchange or personnel changes. They also include information about who the readers are. 

“They are trying to help the models understand what these readers might care about, where their perspectives lie, the relevant communities, and also communities that we actually would not want to include in this newsletter, as well as the relevant school districts,” he said. “We’re passing all this information into Scrape now.”

Editors are pleased with the change, Kevin said. They are now refining the descriptions to produce more actionable tips.

The result: Scrape evolved from a small experiment to “load-bearing and critical infrastructure” for six (and soon, eight) newsletters, with nearly 20 journalists across the newsroom following its output.  The editors’ skills and audience demand has made the local newsletters a success. Scrape is helping the Local team meet demand without overworking the desk. 

What we’d do differently 

One of the hardest problems was generalizing from one editor and geography to many. The original design philosophy was region-based: scrape everything about, say, Lower Merion, then filter using a single “newsworthiness” prompt. That worked “manageably” for one editor and a few newsletters, but did not scale, Kevin said.

Instead of treating “newsworthiness” as a generic, abstract concept, they now anchor it in specific newsletter briefs that editors author themselves – who the readers are, what they’re interested in, and where the coverage line stops.

Another lesson involves cost and search strategy. Early on, the team tried to balance quality with the high cost of searching many small, scattered sources. They discovered that prompting can implicitly control search depth and behavior, but only after trial and error.

In hindsight, Kevin suggests being more explicit, earlier, about acceptable cost-quality trade-offs and designing prompts and architecture around those constraints from day one. 

Here’s how you can apply this 

Scrape’s journey offers a practical blueprint for newsrooms and organizations trying to use AI for finding news in messy local data.

  • Start with a real workflow, not an abstract idea. Identify a specific person or team who is already doing the work manually and measure their effort (e.g., “12 hours a week just to find news”).
  • Turn existing outputs into training data. Use their current products (tip sheets, newsletters, etc.) as both a target and a benchmark.
  • Create a tight human-in-the-loop feedback loop. Ask editors to annotate outputs daily: what’s good, what’s noise, and what should never appear (examples from this case include routine road closures or recurring fitness-class promotions).
  • Encode editorial judgment explicitly. Replace vague notions of “newsworthiness” with concrete prompts and newsletter descriptions: audience, geography in and out of scope, recurring beats, and examples of past coverage.
  • Design for cost and quality together. Decide up front how much search depth you can afford, and use prompting and architecture to stay within that envelope.

The Lenfest AI Collaborative and Fellowship Program is supported by OpenAI and Microsoft. 

联系我们 contact @ memedata.com