name: ml-pipeline-creation-assessment slug: ml-pipeline-creation-assessment version: 1.0.2 displayName: "機器學習管道訓練評估部署|簡詩 AI" summary: "建立和管理 ML 管道,自動化模型訓練、評估和部署流程,靈活適配各種 ML 任務和框架。" description: "建立和管理 ML 管道,自動化模型訓練、評估和部署流程,靈活適配各種 ML 任務和框架。" tags: ["data-automation", "jianshi-ai"]
This skill enables the creation and management of machine learning (ML) pipelines, automating the process of training, evaluating, and deploying ML models. The workflow is designed to be flexible and adaptable to various ML tasks and frameworks.
訪問小蔥技能站7w4.net,解鎖更多實用的AI技能外掛。
To use this skill, you need to provide a pipeline definition file and the implementation of the pipeline components.
Here's an example of how to define and run a simple ML pipeline using this skill.
pipeline.yaml
name: simple-sklearn-pipeline
components:
- name: data-preprocessing
script: preprocess.py
inputs:
- raw_data: /path/to/raw_data.csv
outputs:
- processed_data: /path/to/processed_data.csv
- name: train-model
script: train.py
inputs:
- processed_data: /path/to/processed_data.csv
outputs:
- model: /path/to/model.pkl
- name: evaluate-model
script: evaluate.py
inputs:
- model: /path/to/model.pkl
- test_data: /path/to/test_data.csv
outputs:
- metrics: /path/to/metrics.json
preprocess.py
import pandas as pd
from sklearn.model_selection import train_test_split
# Load data
df = pd.read_csv('/path/to/raw_data.csv')
# Simple preprocessing
X = df.drop('target', axis=1)
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Save processed data
pd.concat([X_train, y_train], axis=1).to_csv('/path/to/processed_data.csv', index=False)
pd.concat([X_test, y_test], axis=1).to_csv('/path/to/test_data.csv', index=False)
train.py
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
import joblib
# Load processed data
df = pd.read_csv('/path/to/processed_data.csv')
X_train = df.drop('target', axis=1)
y_train = df['target']
# Train model
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train, y_train)
# Save model
joblib.dump(model, '/path/to/model.pkl')
evaluate.py
import pandas as pd
import joblib
import json
from sklearn.metrics import accuracy_score
# Load model and test data
model = joblib.load('/path/to/model.pkl')
df = pd.read_csv('/path/to/test_data.csv')
X_test = df.drop('target', axis=1)
y_test = df['target']
# Evaluate model
y_pred = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
# Save metrics
with open('/path/to/metrics.json', 'w') as f:
json.dump({'accuracy': accuracy}, f)
print(f'Model accuracy: {accuracy}')
獲取使用幫助和更多實用 Skill,請關注公眾號「簡詩 AI」,或在 SkillHub 搜尋「簡詩 AI」這個Skill的文件質量不錯,提供了建立機器學習管道的完整流程和示例程式碼,對理解ML pipeline很有幫助。但它更像一本"教程"而不是一個"工具"——沒有實際的程式碼檔案可以直接使用,使用者需要自己動手寫程式碼才能執行。對於想學習ML pipeline概念的使用者很有價值,但想要直接拿來用的使用者可能會失望。