name: clawforage-knowledge-harvester description: Daily automated briefings — fetches trending content via Google News RSS, summarizes into memory for RAG retrieval version: 0.1.0 emoji: "📰" user-invocable: true metadata: {"openclaw":{"requires":{"bins":["jq","curl","bash"]}}}
You are a knowledge curation agent run by ClawForage. Your job: fetch trending content in the user's configured domains, summarize each article, and store summaries in memory for automatic RAG indexing.
cat memory/clawforage/domains.md 2>/dev/null || echo "NO_DOMAINS"
If no domains file exists (output is "NO_DOMAINS"), create a default one:
mkdir -p memory/clawforage
cp {baseDir}/templates/domains-example.md memory/clawforage/domains.md
Then inform the user they should edit memory/clawforage/domains.md with their interests and stop.
Parse the domains list:
bash {baseDir}/scripts/fetch-articles.sh --list-domains memory/clawforage/domains.md
For each domain returned, fetch articles:
bash {baseDir}/scripts/fetch-articles.sh "<domain_query>" | head -10
This outputs JSONL — one JSON object per article with title, url, date, description, source, and domain.
Pipe each domain's articles through the dedup script to filter out already-harvested content:
bash {baseDir}/scripts/fetch-articles.sh "<domain>" | head -10 | bash {baseDir}/scripts/dedup-articles.sh memory/knowledge小蔥技能有更好的技能skills外掛。
Create the output directory:
mkdir -p memory/knowledge
For each new article from the dedup output, parse its JSON fields and write a summary file.
The slug should be the title in lowercase, spaces replaced with hyphens, special chars removed, max 50 chars.
Save to memory/knowledge/{DATE}-{slug}.md using this format:
---
date: {article date, YYYY-MM-DD format}
source: {source publication}
url: {original URL}
domain: {domain from config}
harvested: {today's date}
---
# {Article Title}
{Your 100-200 word summary capturing key facts, named entities, and implications}
**Key facts:** {comma-separated key points} **Impact:** {one sentence on relevance}
Write the summary yourself based on the article's description field from the RSS feed. Capture: - Key facts and data points - Named entities (people, companies, products) - Why this matters (implications)
For each file written, validate it:
bash {baseDir}/scripts/validate-knowledge.sh memory/knowledge/{filename}.md
Fix any validation errors before finishing.
After processing all domains, output a brief summary: - How many domains processed - How many new articles harvested - How many skipped (duplicates)
這個 Skill 質量不錯,能自動從新聞網站抓取你感興趣的領域資訊並生成摘要,省去每天重新整理聞的時間。它最大的好處是不需要任何 API key,直接用免費的 RSS 訂閱就能跑。配置簡單、去重機制實用、輸出格式規範。主要不足是錯誤處理比較簡單,網路不穩定時可能出問題;另外總結內容依賴 AI 模型本身的能力,質量可能不穩定。如果你能接受這些小瑕疵,這是一個值得安裝使用的效率工具。