NLP · 2026 — present · Active

Pantip SET Sentiment Monitor

An automated NLP pipeline that scrapes Thai investor discussions from Pantip.com, links them to SET-listed tickers, scores sentiment with XLM-RoBERTa and serves a live Streamlit dashboard.

Sentiment trend chart with rolling sentiment lines for several SET tickers and dashed closing-price overlays
Multi-ticker rolling-average sentiment with a dual-axis closing-price overlay from yfinance.

Summary

Problem
Thai retail investors discuss SET stocks every day on Pantip.com, Thailand's largest public web forum — questions like "Should I buy DELTA right now?" — but that signal is unstructured, informal and never linked to the stocks it is about.
Solution
A scheduled pipeline on free-tier infrastructure — Selenium scraper, 4-pass entity matcher, multilingual XLM-RoBERTa sentiment scoring, Turso database, Streamlit dashboard and a versioned Kaggle dataset — running end to end every 3 hours via GitHub Actions.
Result
A lag-correlation backtest (lags 0–5 days) found no consistent signal from Pantip sentiment to SET price moves — the project's main research finding. The published dataset scores 10.00/10.00 for Kaggle usability.
Kaggle usability
10.00 / 10.00
Entity matching
4 passes
symbol, company name, alias, fuzzy
Confidence cutoff
≥ 0.65
about 2× the 3-class random baseline
Refresh
every 3 h
GitHub Actions schedule
Backtest lags
0–5 days

Project Summary

Pantip SET Sentiment Monitor is an automated NLP pipeline that mines Thai retail investor discussions from Pantip.com — Thailand’s largest public web forum — and scores them for sentiment against SET-listed stocks. Thai retail investors post there daily: questions like “Should I buy DELTA right now?” or opinions like “PTT is looking strong this week.” This project captures that signal, links each post to the stock it is talking about, and quantifies the mood.

The system runs entirely on free-tier infrastructure — GitHub Actions for scheduling, Turso (LibSQL) as a cloud database, and Streamlit Community Cloud for the dashboard — and refreshes automatically every 3 hours.

Automation Pipeline

Each scheduled run executes the full chain end to end: scrape → entity match → NLP inference → database write → Kaggle export. A local SQLite fallback is used during development when the Turso cloud database is not configured.

Pipeline diagram from Pantip.com through posts, ticker links and scores to alerts, Kaggle export and the dashboard

Data flow of the automated pipeline: Pantip.com → posts → post-ticker links → scores → alerts, Kaggle export and dashboard.

Collection. A Selenium-based scraper visits Pantip’s five investment tag pages and extracts post titles, reply counts, and timestamps. Because Pantip renders both timestamps and reply counts via JavaScript rather than raw HTML, a separate AJAX endpoint was reverse-engineered to recover accurate reply counts, and a multi-strategy timestamp parser handles Open Graph tags, JSON-LD, and data attributes. A circuit breaker and randomised delays between requests handle page-load failures and rate limits gracefully.

Entity matching. A 4-pass linker runs on each post body: (1) exact ticker symbol match, (2) full Thai company name match, (3) alias dictionary lookup covering about 50 hand-curated Thai nicknames, and (4) RapidFuzz approximate matching on PyThaiNLP-tokenised text at a threshold of 92. Each match records the method used and a confidence score. Aliases shorter than 8 characters are excluded from fuzzy matching to prevent short strings matching inside unrelated long posts.

Sentiment scoring. Posts are scored with cardiffnlp/twitter-xlm-roberta-base-sentiment — a multilingual XLM-RoBERTa model fine-tuned on Twitter data across 30+ languages including Thai. Each post body is tokenised and passed through the model in batches; the winning class becomes the label (positive / neutral / negative) and its softmax probability the confidence score. Only predictions with confidence ≥ 0.65 are stored — roughly 2× the random-chance baseline of 0.33 for a 3-class model.

Storage and analysis. Results are written to a Turso (LibSQL) cloud database across 6 normalised tables: posts, post_tickers, scores, alerts, prices, and kaggle_exports. A z-score (≥ 2.5σ) and volume-surge (3× the 7-day average) anomaly detector runs after scoring to flag statistically significant sentiment spikes. A Pearson/Spearman lag-correlation engine then fetches daily closing prices from yfinance and computes sentiment-to-price lag correlations for each ticker.

Dashboard with a ranked horizontal bar chart of ticker sentiment and summary cards

Sentiment dashboard: ranked sentiment for every tracked SET ticker, daily post volume and KPI summary cards.

Analysis & Findings

The pipeline conversion funnel shows that most data loss happens before entity linking — not during NLP scoring. Posts with no body text or no mention of a tracked ticker are dropped at the linking stage and never reach the model, so the scored dataset is a filtered subset of scraped posts, not a complete record of all Pantip investment discussion.

Funnel from scraped posts to posts linked to a ticker to posts scored by the model

Post-to-score conversion funnel: most loss occurs before linking, from posts with no body text or no tracked ticker.

Entity matching is dominated by exact symbol matches, which pile up at 1.0 confidence because the ticker appeared literally in the post text. Fuzzy matches account for roughly 2% of all links and carry lower, more spread-out confidence scores — the alias dictionary and fuzzy pass catch genuine mentions that exact matching would miss, but contribute a small fraction of the total.

Horizontal bar chart comparing exact, alias and fuzzy matches

Entity-matching quality: all ticker links by method (exact, alias, fuzzy) against the subset that reached NLP scoring.

Model confidence shows minimal difference between positive, neutral, and negative predictions. That reflects genuine ambiguity in short, informal Thai text rather than poor performance; for the most reliable analysis, filtering to confidence ≥ 0.85 is recommended. Approximately 85% of posts are classified as neutral. This is expected — Pantip’s investment boards are predominantly Q&A style, where most threads ask questions rather than express strong directional opinions — and the posts with positive or negative labels carry the analytical signal.

Confidence histogram and a bar chart of positive, neutral and negative labels

Model confidence histogram by predicted label and the sentiment label distribution, with the neutral majority expected for a Q&A forum.

Ticker concentration analysis using the Herfindahl-Hirschman Index (HHI) shows that a small number of high-profile tickers — particularly large-cap stocks like DELTA, PTT, and KBANK — account for a disproportionate share of discussion volume. Sentiment for tickers with fewer than 10 posts should be treated with caution, as a single strong opinion can skew the average.

Horizontal bar chart of each ticker's share of post volume

Ticker concentration (HHI): which SET stocks dominate Pantip discussion and how evenly post volume is distributed.

The lag-correlation backtest across lags 0–5 days between daily sentiment aggregates and yfinance closing-price returns found no consistent predictive signal for any ticker. Correlation coefficients clustered near zero across all lags, suggesting that Pantip retail sentiment in its current form does not reliably lead SET price movement. This is the main research finding of the project.

Bar chart of correlation coefficients for lags 0 to 5 days

Pearson/Spearman lag-correlation backtest across lags 0–5 days between daily sentiment and closing-price returns.

Engagement analysis found no meaningful relationship between reply count and sentiment direction or intensity (Pearson r ≈ 0). Posts on Pantip do not attract more comments simply because they are more negative or more emotionally charged — discussion volume appears driven by topic relevance rather than sentiment polarity.

Two scatter plots of reply count versus sentiment

Engagement correlation: reply count against sentiment score and intensity.

Kaggle Dataset & Notebook

The scored dataset is exported to Kaggle on every pipeline run as a versioned, dated CSV. Each export contains 12 columns: post_id, title_th, url, replies, posted_at, ticker, match_confidence, match_method, sentiment, confidence, label, and scored_at. The dataset is published under CC BY-SA 4.0 and achieved a 10.00/10.00 Kaggle usability score.

A beginner-friendly EDA notebook is published alongside the dataset with 9 guided exercises across three difficulty levels: data loading and date handling, entity-matching quality, model confidence, ticker sentiment rankings, time-series trends, engagement correlation, day-of-week patterns, a ticker-by-month sentiment heatmap, and a lag-1 sentiment-to-volume prediction test.

Kaggle dataset page for pantip-set-sentiment with usability score and file preview

The Pantip SET Sentiment dataset on Kaggle.

Resources

Dataset
Pantip SET Sentiment · CC BY-SA 4.0 (opens in a new tab)Versioned CSV exported on every pipeline run, with a companion EDA notebook.