ML-Project
xenodeve · ส่วนตัว
Final Project for Machine Learning
สรุปโปรเจกต์
ภาพรวม
โปรเจกต์นี้คือระบบประเมินและเปรียบเทียบประสิทธิภาพของอัลกอริทึม Machine Learning โดยทดสอบ 4 วิธีการแบ่งข้อมูล (Holdout, Random Subsampling, Repeated CV, Bootstrap) บนชุดข้อมูลมะเร็งเต้านม Wisconsin
สรุป 30 วินาที
- ประเมิน 2 ตัวจำแนกประเภท (Logistic Regression, Linear SVM) พร้อมรายงานค่าสถิติและช่วงความเชื่อมั่น 95% - จัดการด้วย Jupyter Notebook ที่รันซ้ำได้ พร้อมภาพสรุปประสิทธิภาพอัตโนมัติ
เหมาะกับใคร
เหมาะสำหรับนักพัฒนา ML, นักวิเคราะห์ข้อมูล และทีมวิจัยที่ต้องการกรอบการตรวจสอบประสิทธิภาพโมเดลที่โปร่งใสและทำซ้ำได้
- หมวดหมู่
- Machine Learning
- ปี
- 2026
- ผู้ดูแลผลงาน
- xenodeve
รายละเอียดเพิ่มเติมรายละเอียดเชิงลึก
โปรเจกต์นี้มุ่งเน้นการประเมินและเปรียบเทียบวิธีการตรวจสอบความถูกต้องของโมเดลแมชชีนเลิร์นนิง 4 รูปแบบ ได้แก่ Holdout, Random Subsampling, Repeated Stratified K-Fold และ Bootstrap โดยทดสอบกับชุดข้อมูลมะเร็งเต้านมวิสคอนซินจำนวน 569 ตัวอย่าง
การวิเคราะห์ครอบคลุมตัวจำแนกประเภท 2 แบบ คือ Logistic Regression และ Linear SVM พร้อมคำนวณค่าความแม่นยำ ความแม่นยำกลับคืน ค่า F1-score รวมถึงค่าเฉลี่ย ส่วนเบี่ยงเบนมาตรฐาน และช่วงความเชื่อมั่น 95% เพื่อให้เห็นภาพรวมของประสิทธิภาพโมเดลอย่างชัดเจน
ผลลัพธ์ถูกจัดระเบียบในรูปแบบ Jupyter Notebook พร้อมไฟล์ CSV สำหรับข้อมูลสถิติและภาพกราฟิกแบบ Bar chart และ Boxplot ที่ช่วยในการตีความผลลัพธ์เชิงลึกอย่างมีระบบ
เอกสารต้นฉบับรายละเอียด (README)
Model Evaluation Benchmark
Owner Name: Theerut Dokkatin
Student ID: 6604022630250
Project Goal
Implement and compare four model evaluation methods on a real medical classification dataset.
Course Requirement Coverage
- 4 methods implemented: Holdout, Random Subsampling, Repeated CV, Bootstrap
- At least 2 classifiers used
- Reported mean, std, 95% CI for each method-classifier pair
- Includes bar chart and boxplot visualizations
- Notebook contains code + output cells
Dataset
- Name: Breast Cancer Wisconsin
- Source: sklearn.datasets.load_breast_cancer
- Samples: 569
- Features: 30
- Task: Binary classification
Methods Implemented
- Holdout (stratified, 10 iterations with different seeds)
- Random Subsampling (
R=60) - Repeated Stratified K-Fold (
10x5) - Bootstrap OOB (
B=300)
Classifiers
- Logistic Regression
- Linear SVM
Reported Metrics
- Accuracy
- Precision (macro)
- Recall (macro)
- F1-score (macro)
- Mean, Std, 95% CI for each method-classifier combination
Outputs
- Notebook: 6604022630250_Project.ipynb
- Summary CSV: outputs/metrics_summary.csv
- Long-format scores: outputs/all_scores_long.csv
- Figures:
- figures/holdout_confusion_matrices.png
- figures/benchmark_bar_boxplot.png
How To Run
- Open notebook: 6604022630250_Project.ipynb
- Run all cells from top to bottom
- Confirm these files are generated:
- outputs/metrics_summary.csv
- outputs/all_scores_long.csv
- figures/holdout_confusion_matrices.png
- figures/benchmark_bar_boxplot.png
Submission Checklist
- Jupyter Notebook (.ipynb) with all code and output cells
- PDF export of the notebook
- All figures saved as PNG (dpi >= 150)
- README.md
- ZIP package named with student ID (example: 6604022630250_Project.zip)
Notes
- Test data is not used for training in Holdout.
- Repeated stratified splits are used to reduce variance.
- Pipelines are used to prevent data leakage.
Writing Style
- Core language: English
- Thai insight: ใช้ภาษาไทยเสริมเพื่ออธิบายเหตุผลเชิงตีความและข้อสังเกตสำคัญ
ถาม AI เกี่ยวกับ ML-Project
ถามแนวทาง เทคโนโลยี หรือความเหมาะสมของโปรเจกต์นี้กับไอเดียของคุณ
กำลังโหลดแชท / Loading chat…