Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Postest-AI_TPLUS

Total questions: 5

Worksheet time: 10mins

Name
Class
Date
1.

ในระบบ NOC ของ TPLUS Digital มี alarm logs 10,000 records โดยมี anomaly จริงอยู่ประมาณ 2% (200 records) หากใช้ Isolation Forest แล้วตั้งค่า contamination=0.05 ผลที่จะเกิดขึ้นคืออะไร และส่งผลต่อการดูแลเครือข่ายอย่างไร?

a)

Model จะ label ~200 records เป็น anomaly (2%) เพราะรู้จำนวนจริงอยู่แล้ว

b)

Model จะ label ~500 records เป็น anomaly (5%) ทำให้ False Positive สูง NOC Engineer ต้องเสียเวลาตรวจสอบ alarm ที่ไม่ใช่ปัญหาจริงมากขึ้น

c)

Contamination ไม่มีผลต่อ Isolation Forest เพราะใช้แค่ path length

d)

Model จะ error เพราะ contamination ต้องเท่ากับสัดส่วน anomaly จริงเสมอ

2.

Complaint log ของ TPLUS Digital มี 2 ข้อความดังนี้: Doc1: 'internet slow internet problem' Doc2: 'internet down' เมื่อใช้ TfidfVectorizer (sklearn default: smooth_idf=True) คำว่า 'internet' จะมีค่า TF-IDF ใน Doc1 VS Doc2 อย่างไร?

a)

Doc1 สูงกว่า เพราะ TF ของ 'internet' ใน Doc1 = 2/4 > Doc2 = 1/2

b)

เท่ากัน เพราะ IDF เท่ากัน และ TF ถูก normalize แล้ว

c)

Doc2 สูงกว่า เพราะเอกสารสั้นกว่า ความหนาแน่นคำสูงกว่า

d)

ทั้งสองเอกสาร: TF-IDF ของ 'internet' ต่ำมาก เพราะ 'internet' ปรากฏในทุก doc → IDF ใกล้ 0 (ไม่ discriminative)

3.

เดล Auto-routing ของ TPLUS Digital ให้ผลดังนี้:

Category | Precision | Recall

-----------|-----------|-------

Network | 0.90 | 0.95

Billing | 0.85 | 0.50

SIM | 0.70 | 0.80

Other | 0.60 | 0.75

Category ใดที่ควรแก้ไขเร่งด่วนที่สุด และเพราะอะไร?

a)

Other: ควรแก้ก่อนเพราะ Precision ต่ำสุด (0.60)

b)

Billing: ควรแก้ก่อนเพราะ Recall ต่ำสุด (0.50) = โมเดลส่ง Billing complaints ไปผิด dept ถึง 50% ลูกค้าไม่ได้รับการแก้ไขปัญหา → churn risk สูง

c)

SIM: ควรแก้ก่อนเพราะ Precision ต่ำ (0.70) ทำให้ SIM team ทำงานซ้ำซ้อน

d)

Network: ควรแก้ก่อนเพราะมีผู้ใช้ร้องเรียน Network มากที่สุด

4.

Dashboard ของ TPLUS Digital โหลด ML model ขนาด 120MB ทุกครั้งที่ผู้ใช้กด predict ปุ่ม → หน้าจอค้างนาน 8 วินาที Code ปัจจุบัน:

def load_model()

: return joblib.load('model.pkl')

model = load_model() # เรียกทุกครั้งที่ rerun

วิธีใดแก้ปัญหาได้ถูกต้องที่สุดใน Streamlit?

a)

ใช้ @st.cache_data เพราะ cache ข้อมูลได้ทุกประเภท

b)

ใช้ @st.cache_resource เพราะออกแบบมาสำหรับ shared resource เช่น ML model, DB connection ที่ไม่ควร serialize ซ้ำ

c)

ใช้ st.session_state['model'] เพื่อเก็บ model ต่อ session

d)

ย้าย load_model() ไปไว้ใน global scope ก่อน st.title()

5.

TPLUS Digital ต้องการ predict 'Major Network Incident'

ล่วงหน้า 2 ชั่วโมง โดยใช้ KPI รายชั่วโมง:

Drop Call Rate, RSSI, Throughput, Alarm Count

ข้อมูล 1 ปี = 8,760 rows

Major Incident เกิดขึ้น 87 ครั้ง (1% ของทั้งหมด)

นักเรียนเทรน Logistic Regression ได้ Accuracy = 99.1%

แต่ทีม NOC บอกว่า model ไม่มีประโยชน์เลย

เพราะอะไร และควรแก้ไขอย่างไร?

a)

99.1% ต่ำเกินไป ต้องเพิ่ม features มากขึ้นจึงจะดี

b)

Model overfit เพราะ training data มีแค่ 1 ปี ต้องเก็บข้อมูลเพิ่ม

c)

Class Imbalance: Model เรียนรู้ว่า 'predict Normal ตลอด' ก็ได้ 99% accuracy จริงๆ แล้ว model ไม่ detect incident ได้เลยสักครั้ง ควรดู F1-Score / Recall ของ class Incident แทน

d)

Logistic Regression ไม่เหมาะกับ time-series ต้องใช้ LSTM แทน