ZYVOPMulti-Platform Sync
SeriesAI NewsWhy ZyVOPJoin Discord
LoginGet Started
ZYVOPMulti-Platform Sync

The Developer Publishing Hub. Write once, publish everywhere, and make your work citation-ready with built-in SEO, AEO, and GEO discovery support. Zero reader paywalls.

Content

  • Categories
  • Tags
  • Badges
  • Leaderboard
  • Write Article
  • Newsletter

Company

  • About Us
  • Why ZyVOP
  • Changelog
  • Compare Platforms
  • Hashnode vs ZyVOP
  • DEV vs ZyVOP
  • Developer API & CLI
  • Author Handbook
  • Contact

Connect

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • DMCA Policy
  • Code of Conduct

© 2026 ZyVOP. Developer Publishing Hub.

Zero paywalls · Full content ownership
All systems operational
HomeAI Reads Vietnamese Slang as Angry. 3 Models Don't.

AI Reads Vietnamese Slang as Angry. 3 Models Don't.

Lê Đức Minh
Lê Đức Minh
AI Engineer
September 9, 2026
6 min read
AI Reads Vietnamese Slang as Angry. 3 Models Don't.
Article
👍2

Almost every paper on Vietnamese sentiment analysis reports impressive accuracy. 94% F1 is the standard figure everyone cites.

That number is real, but it usually comes from clean academic datasets where students write polite course feedback.

What happens when your model encounters real comments from Vietnamese social media, full of slang, sarcasm, and friends calling each other "tao" and "mày" — the same messy real-world text that made me start benchmarking Vietnamese BERT models in the first place?

I tested four modern LLMs on both types of text. On curated sentences, all four models scored between 84% and 88%. On real social media comments, three models barely flinched.

Llama-3.1-8B-Instruct dropped 20 points, repeatedly misinterpreting friendly banter as outright hostility.

The Test Setup

I used two public datasets:

  1. UIT-VSFC: Formal student feedback. Clean, structured sentences.

  2. UIT-VSMEC (Relabelled): Real Facebook comments featuring slang, missing diacritics, and informal spelling.

Both datasets were normalized to a standard 3-class sentiment scheme (positive, negative, neutral). I sampled 25 random items from each under zero-shot prompting at temperature 0.

Model

Curated Feedback (VSFC)

Social Media (VSMEC)

Accuracy Drop

Qwen3-8B

84% (21/25)

80% (20/25)

4 pts

Llama-3.1-8B-Instruct

88% (22/25)

68% (17/25)

20 pts

gpt-4o-mini

84% (21/25)

80% (20/25)

4 pts

DeepSeek-V4-Flash

84% (21/25)

76% (19/25)

8 pts

Three models stayed within a tight 4 to 8 point drop. Llama-3.1 was the clear outlier, dropping more than double its peers.

The Anatomy of the Misclassification

Digging into the individual errors revealed a consistent pattern.

Here are four positive comments from the test set:

Comment Text

Meaning

Qwen3-8B

Llama-3.1

gpt-4o-mini

DeepSeek-V4

"...nghe hay hơn bản gốc...nhiều < 3"

Praising audio quality, ending with a heart

Positive

Negative

Positive

Positive

"con gái tao thì suốt ngày hêy siri bắt chước mẹ 😂"

Fondly describing daughter with a laugh emoji

Positive

Negative

Positive

Positive

"per hẹn xem phim này nữa nha mày 😛"

Friendly invite to a movie with playful emoji

Positive

Negative

Positive

Positive

"nghe bạn này nói dễ thương zị"

"This person speaks so cutely" (slang spelling)

Positive

Negative

Positive

Positive

Three models recognized warmth through slang and emojis.

Llama-3.1 got all four wrong.

Two comments used "tao" and "mày". In formal Vietnamese, these pronouns can sound abrasive. Between close friends online, they are completely ordinary.

The other two examples contained no rough pronouns, just informal spelling ("zị") and emoticons ("< 3"). Llama seems calibrated to assume that any informal or non-standard Vietnamese text is inherently negative.

This is a classic silent failure: the model delivers its verdict with absolute confidence, giving your backend no indication that it misunderstood the conversational register — exactly the kind of invisible failure that makes the hard parts of being an AI engineer hard.

def call_model(model, text):
    prompt = (
        f'Phân loại cảm xúc của câu sau là "positive", "negative", hoặc "neutral". '
        f'Chỉ trả lời đúng một từ, không giải thích.\n\nCâu: "{text}"'
    )
    resp = requests.post(
        url, headers={"Authorization": f"Bearer {token}", "Content-Type": "application/json"},
        json={"model": model, "messages": [{"role": "user", "content": prompt}],
              "max_tokens": 800, "temperature": 0.0},
        timeout=60,
    )
    resp.raise_for_status()
    msg = resp.json()["choices"][0]["message"]
    return (msg.get("content") or "").strip().lower()

A Practical Note on Token Budgets

Notice max_tokens=800 in the script above.

When set to 300 tokens, Qwen3-8B frequently failed. Because it uses internal reasoning, it exhausted 300 tokens "thinking" about a one-word label and timed out before emitting the answer. Raising the limit to 800 resolved the issue.

If you are running reasoning models on simple classification tasks, verify that your max token limits leave room for their internal chain of thought.

Experiment

Here's the actual run, step by step. The complete script is in the appendix at the end of this post.

1. Sample 25 items from each dataset with a fixed seed, mapping VSMEC's numeric labels onto the same 3-class scheme as VSFC:

VSMEC_LABEL_MAP = {-1: "negative", 0: "neutral", 1: "positive"}

def sample_vsfc(n):
    ds = load_dataset("ura-hcmut/UIT-VSFC")["test"]
    idx = list(range(len(ds)))
    random.Random(SEED).shuffle(idx)
    picked = idx[:n]
    return [{"text": ds[i]["text"], "gold": ds[i]["label"]} for i in picked]

def sample_vsmec(n):
    ds = load_dataset("viethq1906/UIT-VSMEC-Sentiment-Relabelled")["test"]
    idx = list(range(len(ds)))
    random.Random(SEED + 1).shuffle(idx)
    picked = idx[:n]
    return [{"text": ds[i]["sentence"], "gold": VSMEC_LABEL_MAP[ds[i]["sentiment"]]} for i in picked]

2. Parse the model's raw output into one of the three labels — or None if it doesn't say any of them:

def extract_label(raw):
    raw = raw.lower()
    for label in ("positive", "negative", "neutral"):
        if label in raw:
            return label
    return None

3. Run every model against every sampled comment, in both datasets:

for dataset_name, samples in [("vsfc", vsfc_samples), ("vsmec", vsmec_samples)]:
    for ex in samples:
        entry = {"text": ex["text"], "gold": ex["gold"], "models": {}}
        for model in MODELS:
            try:
                raw, finish = call_model(model, ex["text"])
                label = extract_label(raw)
                entry["models"][model] = {"raw": raw, "extracted": label, "finish": finish,
                                           "correct": label == ex["gold"]}
            except Exception as e:
                entry["models"][model] = {"error": str(e)[:200]}
        results[dataset_name].append(entry)

4. Compute accuracy per model, per dataset — this is the table at the top of the post:

summary = {}
for dataset_name in ("vsfc", "vsmec"):
    for model in MODELS:
        correct = sum(1 for e in results[dataset_name] if e["models"].get(model, {}).get("correct"))
        total = len(results[dataset_name])
        summary.setdefault(model, {})[dataset_name] = f"{correct}/{total} ({100*correct/total:.0f}%)"

Run it yourself: uv run python sentiment_gap_run.py.

Takeaways

  • Curated scores don't reflect social listening reality. High benchmark numbers on formal feedback don't guarantee resilience to online slang.

  • Other models handle Vietnamese slang gracefully. Qwen3, gpt-4o-mini, and DeepSeek-V4 held 76–80% accuracy on real comments.

  • Model selection also matters as prompt tweaks. If your pipeline processes social media text, benchmark candidate models on real slang before deploying.

References

  • UIT-VSFC — the curated student-feedback dataset.

  • UIT-VSMEC (Relabelled) — the real social-media comment dataset.

  • Llama-3.1-8B-Instruct — the model that misread slang as hostile.

  • Qwen3-8B — one of the models that held steady.

  • gpt-4o-mini — OpenAI's model documentation.


Have you noticed LLMs misinterpreting informal language in your domain? Share your findings below.

👉 Follow my work: LinkedIn | GitHub

Appendix: Full Script

For anyone who wants the complete, runnable file:

#!/usr/bin/env python3
"""Measure whether LLM-prompted sentiment classification holds up on real Vietnamese
social media text (UIT-VSMEC) the way it does on curated, formal text (UIT-VSFC).

Both datasets are public, real, human-labeled:
- ura-hcmut/UIT-VSFC (test split, 3166 rows) — formal student feedback, 3-class
  (positive/negative/neutral).
- viethq1906/UIT-VSMEC-Sentiment-Relabelled (test split, 693 rows) — real Facebook
  comments, slang/emoji/typos, sentiment relabelled to the same 3-class scheme
  (-1/0/1 = negative/neutral/positive).

No fine-tuning here: this tests LLM-prompted classification specifically, since a lot
of 2026 production sentiment analysis is done via LLM prompting rather than a
dedicated fine-tuned classifier. Not a reproduction of the older PhoBERT/ensemble
benchmark numbers (94% VSFC / ~60% VSMEC CNN baseline) cited in prior literature —
those are a different method entirely, cited separately in the post as corroboration.
"""
import json
import os
import random
import time
from pathlib import Path

import requests
from datasets import load_dataset

OUT_PATH = Path("content/2026-09-01/sentiment-gap/scratch/sentiment_gap_results.json")
N_PER_DATASET = 25
SEED = 20260901

HF_MODELS = [
    "Qwen/Qwen3-8B",
    "meta-llama/Llama-3.1-8B-Instruct",
]
OPENROUTER_MODELS = [
    "openai/gpt-4o-mini",
    "deepseek/deepseek-v4-flash-0731",
]
MODELS = HF_MODELS + OPENROUTER_MODELS

HF_TOKEN = os.environ["HF_TOKEN"]
OPENROUTER_API_KEY = os.environ.get("OPENROUTER_API_KEY")
HF_ROUTER_URL = "https://router.huggingface.co/v1/chat/completions"
OPENROUTER_URL = "https://openrouter.ai/api/v1/chat/completions"

VSMEC_LABEL_MAP = {-1: "negative", 0: "neutral", 1: "positive"}


def sample_vsfc(n):
    ds = load_dataset("ura-hcmut/UIT-VSFC")["test"]
    idx = list(range(len(ds)))
    random.Random(SEED).shuffle(idx)
    picked = idx[:n]
    return [{"text": ds[i]["text"], "gold": ds[i]["label"]} for i in picked]


def sample_vsmec(n):
    ds = load_dataset("viethq1906/UIT-VSMEC-Sentiment-Relabelled")["test"]
    idx = list(range(len(ds)))
    random.Random(SEED + 1).shuffle(idx)
    picked = idx[:n]
    return [{"text": ds[i]["sentence"], "gold": VSMEC_LABEL_MAP[ds[i]["sentiment"]]} for i in picked]


def call_model(model, text):
    prompt = (
        f'Phân loại cảm xúc của câu sau là "positive", "negative", hoặc "neutral". '
        f'Chỉ trả lời đúng một từ, không giải thích.\n\nCâu: "{text}"'
    )
    if model in OPENROUTER_MODELS:
        url, token = OPENROUTER_URL, OPENROUTER_API_KEY
    else:
        url, token = HF_ROUTER_URL, HF_TOKEN
    resp = requests.post(
        url,
        headers={"Authorization": f"Bearer {token}", "Content-Type": "application/json"},
        json={"model": model, "messages": [{"role": "user", "content": prompt}],
              "max_tokens": 800, "temperature": 0.0},
        timeout=60,
    )
    resp.raise_for_status()
    data = resp.json()
    msg = data["choices"][0]["message"]
    content = (msg.get("content") or "").strip().lower()
    return content, data["choices"][0].get("finish_reason")


def extract_label(raw):
    raw = raw.lower()
    for label in ("positive", "negative", "neutral"):
        if label in raw:
            return label
    return None


def run():
    vsfc_samples = sample_vsfc(N_PER_DATASET)
    vsmec_samples = sample_vsmec(N_PER_DATASET)
    print(f"Sampled {len(vsfc_samples)} VSFC, {len(vsmec_samples)} VSMEC")

    results = {"vsfc": [], "vsmec": []}
    for dataset_name, samples in [("vsfc", vsfc_samples), ("vsmec", vsmec_samples)]:
        for ex in samples:
            entry = {"text": ex["text"], "gold": ex["gold"], "models": {}}
            for model in MODELS:
                try:
                    raw, finish = call_model(model, ex["text"])
                    label = extract_label(raw)
                    entry["models"][model] = {"raw": raw, "extracted": label, "finish": finish,
                                               "correct": label == ex["gold"]}
                except Exception as e:  # noqa: BLE001
                    entry["models"][model] = {"error": str(e)[:200]}
            results[dataset_name].append(entry)
            print(f"[{dataset_name}] gold={ex['gold']:8s} " +
                  " ".join(f"{m.split('/')[-1]}={entry['models'][m].get('extracted')}" for m in MODELS))

    # accuracy summary
    summary = {}
    for dataset_name in ("vsfc", "vsmec"):
        for model in MODELS:
            correct = sum(1 for e in results[dataset_name] if e["models"].get(model, {}).get("correct"))
            total = len(results[dataset_name])
            summary.setdefault(model, {})[dataset_name] = f"{correct}/{total} ({100*correct/total:.0f}%)"

    print("\n=== Accuracy summary ===")
    for model, d in summary.items():
        print(model, d)

    OUT_PATH.write_text(json.dumps({"results": results, "summary": summary}, ensure_ascii=False, indent=2))
    print(f"\nWrote {OUT_PATH}")


if __name__ == "__main__":
    run()

Comments (0)

Login to post a comment.

Lê Đức Minh
Lê Đức Minh

AI Engineer

Just Another Dev Guy

Subscribe to Lê Đức Minh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.

More from Lê Đức Minh

View profile

The Harness is the Moat: Building a Deterministic Agent Runtime with Context Pruning

Learn how a deterministic harness prunes context, uses a ledger and transactional tool calls to keep LLM agents reliable over many turns.

6 minSep 16

What DeepSeek's Open-Source Agent Harness Gets Right

A few weeks ago, my dad and I spent an afternoon at the kitchen table trying to figure out what an "AI agent" actually is.

8 minSep 16

I Tested 3 LLMs on Vietnamese Idioms. They Got It Backwards.

State‑of‑the‑art LLMs, even Vietnamese‑specialized ones, score below 50 % on the VIVID benchmark of 1,636 authentic Vietnamese idioms and fail completely on a hand‑picked set of six complex idioms. Th

6 minSep 5

The Real Vietnamese LLM Tokenizer Cost (It's Not 4.5x)

Discover why Vietnamese tokenizers cost only 1.05‑2.14× English tokens, not 4.5×, and how updating tokenizers can cut LLM expenses by a third.

7 minSep 5

A Simple Look at TokPress: A Compressor That Uses an LLM's Tokenizer

TokPress compresses tiny JSON log lines by tokenizing with OpenAI's o200k_base tokenizer, then applying LZ77 and rANS for smaller files.

2 minSep 1