Machine Learning

ML-Project

xenodeve · ส่วนตัว

Final Project for Machine Learning

ML-Project
ภาพรวมอ่านจบใน 30 วินาที

สรุปโปรเจกต์

ภาพรวม

โปรเจกต์นี้คือระบบประเมินและเปรียบเทียบประสิทธิภาพของอัลกอริทึม Machine Learning โดยทดสอบ 4 วิธีการแบ่งข้อมูล (Holdout, Random Subsampling, Repeated CV, Bootstrap) บนชุดข้อมูลมะเร็งเต้านม Wisconsin

สรุป 30 วินาที

- ประเมิน 2 ตัวจำแนกประเภท (Logistic Regression, Linear SVM) พร้อมรายงานค่าสถิติและช่วงความเชื่อมั่น 95% - จัดการด้วย Jupyter Notebook ที่รันซ้ำได้ พร้อมภาพสรุปประสิทธิภาพอัตโนมัติ

เหมาะกับใคร

เหมาะสำหรับนักพัฒนา ML, นักวิเคราะห์ข้อมูล และทีมวิจัยที่ต้องการกรอบการตรวจสอบประสิทธิภาพโมเดลที่โปร่งใสและทำซ้ำได้

หมวดหมู่
Machine Learning
ปี
2026
ผู้ดูแลผลงาน
xenodeve
รายละเอียดเพิ่มเติมรายละเอียดเชิงลึก

โปรเจกต์นี้มุ่งเน้นการประเมินและเปรียบเทียบวิธีการตรวจสอบความถูกต้องของโมเดลแมชชีนเลิร์นนิง 4 รูปแบบ ได้แก่ Holdout, Random Subsampling, Repeated Stratified K-Fold และ Bootstrap โดยทดสอบกับชุดข้อมูลมะเร็งเต้านมวิสคอนซินจำนวน 569 ตัวอย่าง

การวิเคราะห์ครอบคลุมตัวจำแนกประเภท 2 แบบ คือ Logistic Regression และ Linear SVM พร้อมคำนวณค่าความแม่นยำ ความแม่นยำกลับคืน ค่า F1-score รวมถึงค่าเฉลี่ย ส่วนเบี่ยงเบนมาตรฐาน และช่วงความเชื่อมั่น 95% เพื่อให้เห็นภาพรวมของประสิทธิภาพโมเดลอย่างชัดเจน

ผลลัพธ์ถูกจัดระเบียบในรูปแบบ Jupyter Notebook พร้อมไฟล์ CSV สำหรับข้อมูลสถิติและภาพกราฟิกแบบ Bar chart และ Boxplot ที่ช่วยในการตีความผลลัพธ์เชิงลึกอย่างมีระบบ

เอกสารต้นฉบับรายละเอียด (README)

Model Evaluation Benchmark

Owner Name: Theerut Dokkatin

Student ID: 6604022630250

Project Goal

Implement and compare four model evaluation methods on a real medical classification dataset.

Course Requirement Coverage

  • 4 methods implemented: Holdout, Random Subsampling, Repeated CV, Bootstrap
  • At least 2 classifiers used
  • Reported mean, std, 95% CI for each method-classifier pair
  • Includes bar chart and boxplot visualizations
  • Notebook contains code + output cells

Dataset

  • Name: Breast Cancer Wisconsin
  • Source: sklearn.datasets.load_breast_cancer
  • Samples: 569
  • Features: 30
  • Task: Binary classification

Methods Implemented

  1. Holdout (stratified, 10 iterations with different seeds)
  2. Random Subsampling (R=60)
  3. Repeated Stratified K-Fold (10x5)
  4. Bootstrap OOB (B=300)

Classifiers

  1. Logistic Regression
  2. Linear SVM

Reported Metrics

  • Accuracy
  • Precision (macro)
  • Recall (macro)
  • F1-score (macro)
  • Mean, Std, 95% CI for each method-classifier combination

Outputs

  • Notebook: 6604022630250_Project.ipynb
  • Summary CSV: outputs/metrics_summary.csv
  • Long-format scores: outputs/all_scores_long.csv
  • Figures:
    • figures/holdout_confusion_matrices.png
    • figures/benchmark_bar_boxplot.png

How To Run

  1. Open notebook: 6604022630250_Project.ipynb
  2. Run all cells from top to bottom
  3. Confirm these files are generated:
    • outputs/metrics_summary.csv
    • outputs/all_scores_long.csv
    • figures/holdout_confusion_matrices.png
    • figures/benchmark_bar_boxplot.png

Submission Checklist

  • Jupyter Notebook (.ipynb) with all code and output cells
  • PDF export of the notebook
  • All figures saved as PNG (dpi >= 150)
  • README.md
  • ZIP package named with student ID (example: 6604022630250_Project.zip)

Notes

  • Test data is not used for training in Holdout.
  • Repeated stratified splits are used to reduce variance.
  • Pipelines are used to prevent data leakage.

Writing Style

  • Core language: English
  • Thai insight: ใช้ภาษาไทยเสริมเพื่ออธิบายเหตุผลเชิงตีความและข้อสังเกตสำคัญ
Project AI

ถาม AI เกี่ยวกับ ML-Project

ถามแนวทาง เทคโนโลยี หรือความเหมาะสมของโปรเจกต์นี้กับไอเดียของคุณ

กำลังโหลดแชท / Loading chat…

ผู้มีส่วนร่วม
คุยกับ AI ดูงานคล้ายกันติดต่อจ้างงาน