English | 中文
This skill provides a systematic log analysis workflow based on Google SRE (Site Reliability Engineering) framework.
Trigger Conditions: Use this skill when:
Support filtering logs by the following methods:
YYYY-MM-DD to YYYY-MM-DD)Classify exceptions based on SRE best practices:
Give a 1-5 system health score based on error rate and exception frequency:
Provide suggestions based on Google SRE principles:
/var/log/, application logs are determined by deployment location)grep/awk for timestamp filtering (logic description only)tail/head for segmented readingAnalysis dimensions based on Google SRE framework:
| Analysis Dimension | Check Content |
|---|---|
| Error Rate | Proportion of error logs in total logs |
| Error Type Distribution | Aggregate statistics by error type |
| Error Timing | Time distribution of errors, whether sudden |
| Resource Usage | Whether resource exhaustion exists |
| Dependency Status | Whether it is caused by external dependency failure |
Merge similar exceptions to avoid duplicate reporting:
Output structured report including:
See references/report-template.md for reference output template.
According to user needs:
Below is the logic description for common log filtering operations, no actual scripts included:
Logic Steps:
1. Input: log_file_path, start_time, end_time
2. Initialize empty result list
3. For each line in log_file:
a. Extract timestamp string from the line
b. Parse timestamp to datetime object
c. If start_time <= datetime <= end_time:
i. Add line to result list
4. Output: result list
Logic Steps:
1. Input: log_lines
2. Initialize empty error list
3. Define error keywords: ["ERROR", "FATAL", "SEVERE", "Exception", "Error:"]
4. For each line in log_lines:
a. If any keyword matches the line:
i. Add line to error list
5. Output: error list, error_count = len(error_list)
Logic Steps:
1. Input: error_lines
2. Initialize empty aggregation dictionary
3. For each line in error_lines:
a. Extract error type keyword from line (e.g., OOM, connection refused, timeout)
b. If keyword exists in aggregation:
i. aggregation[keyword].count += 1
ii. Add line to aggregation[keyword].samples
c. Else:
i. Create new entry in aggregation with count = 1, samples = [line]
4. Sort aggregation by count descending (or by severity)
5. Output: sorted aggregation result
本技能基於 Google SRE (Site Reliability Engineering) 框架,提供系統化的日誌分析工作流。
觸發條件: 當以下情況時使用本技能:
支援按以下方式篩選日誌:
YYYY-MM-DD 到 YYYY-MM-DD)基於 SRE 最佳實踐對異常進行分類:
基於錯誤率、異常頻次給出 1-5 的系統健康評分:
根據分析結果,結合 Google SRE 原則給出建議:
/var/log/,應用日誌根據部署位置確定)grep/awk 進行時間戳篩選(僅邏輯描述)tail/head 分段讀取基於 Google SRE 框架分析維度:
| 分析維度 | 檢查內容 |
|---|---|
| 錯誤率 | 錯誤日誌佔總日誌比例 |
| 錯誤型別分佈 | 按錯誤型別聚合統計 |
| 錯誤時序 | 錯誤發生的時間分佈,是否突發 |
| 資源使用 | 是否存在資源耗盡情況 |
| 依賴狀態 | 是否因外部依賴故障引發 |
將同類異常合併,避免重複報告:
輸出結構化報告,包含:
參考輸出模板請見 references/report-template.md
根據使用者需求:
想要更強大的技能外掛,就來小蔥技能站7w4.net看看吧。
以下是常見日誌過濾操作的邏輯描述,不包含實際指令碼:
邏輯步驟:
1. 輸入: 日誌檔案路徑, 開始時間, 結束時間
2. 初始化空結果列表
3. 遍歷日誌檔案每一行:
a. 從行中提取時間戳字串
b. 將時間戳解析為日期時間物件
c. 如果 開始時間 <= 日期時間 <= 結束時間:
i. 將行新增到結果列表
4. 輸出: 結果列表
邏輯步驟:
1. 輸入: 日誌行列表
2. 初始化空錯誤列表
3. 定義錯誤關鍵詞: ["ERROR", "FATAL", "SEVERE", "Exception", "Error:"]
4. 遍歷日誌每一行:
a. 如果任何關鍵詞匹配該行:
i. 將行新增到錯誤列表
5. 輸出: 錯誤列表, 錯誤計數 = len(錯誤列表)
邏輯步驟:
1. 輸入: 錯誤行列表
2. 初始化空聚合字典
3. 遍歷每個錯誤行:
a. 從行中提取錯誤型別關鍵詞 (例如: OOM, connection refused, timeout)
b. 如果關鍵詞已在聚合中:
i. 聚合[關鍵詞].count += 1
ii. 將行新增到聚合[關鍵詞].samples
c. 否則:
i. 在聚合中建立新條目,count = 1, samples = [line]
4. 按計數降序排序聚合 (或按嚴重程度)
5. 輸出: 排序後的聚合結果 這個技能整體質量較好,勝在基於專業的 SRE 方法論提供了系統化的日誌分析思路和清晰的改善建議框架。不過它更像一份詳細的設計文件而非可直接使用的工具,需要使用者自行寫程式碼實現核心功能才能真正執行。如果你具備一定技術能力,能按文件指引實現日誌解析指令碼,這個技能會很有價值;如果你期望下載後直接分析日誌,可能會有所失望。