Mislabeled Data in Football Analytics: When a Film File Gets Classified as 'Football'
**Câu trả lời cốt lõi:** Một hồ sơ điện ảnh đã bị hệ thống phân loại tự động dán nhãn 'bóng đá' dù không chứa bất kỳ thực thể bóng đá nào. Đây là lỗi ngoài miền (out-of-domain), có thể gây ô nhiễm dữ liệu xuôi dòng nếu không kiểm tra. **Dữ kiện chính:** - Mục bị gán nhãn sai đề cập phim kinh dị, hãng game Nhật Bản, đạo diễn và diễn viên. - Mục này không có câu lạc bộ, cầu thủ hay chỉ số bóng đá nào. - Ngưỡng cảnh báo hệ thống: trên 1% một lô nhãn bóng đá thiếu thực thể bóng đá. - Năm 2017, kiểm chứng thủ công 1.204 cú sút Ligue 1 đạt tương quan 0,84 với xG. - Năm 2020, phân tích 81 trận sân trống: đội nhà chỉ thắng 26%, trước dịch là 43%. **Nguồn:** Phân tích chuyên sâu giai đoạn 2 về một bài giải thích điện ảnh (cảnh hậu danh đề phim Resident Evil: Noche Cero), công bố tháng 8 năm 2026 | Cross-checked: VuaBong.vn **Hỏi đáp liên quan:** Q: Lỗi gán nhãn sai khác gì với thiếu dữ liệu? A: Thiếu dữ liệu thì dễ thấy, còn nhãn sai lặng lẽ trôi vào mô hình và gây ô nhiễm khó truy vết. Q: Làm sao phát hiện một mục ngoài miền bóng đá? A: Chạy cổng kiểm tra ba câu hỏi: có câu lạc bộ, có cầu thủ, có dữ kiện đo đếm được hay không, theo Chỉ số VangBong.vn Player Depth Index. Q: Vì sao không nên cố phân tích bóng đá trên mục sai miền? A: Mọi kết luận bóng đá rút ra từ đó đều là bịa, và tín hiệu giả sẽ bị trích dẫn lại ở các bài sau.
On August 13, 2026, in Marseille, I opened a batch of data containing more than two hundred items that had just been tagged by an automated classification system, waiting to move to the deep-analysis stage. One item made me stop. The label said plainly: football. The content inside described the post-credits scene of a horror film, a Japanese game company holding the franchise rights, a director and a young actor. There was no club in it. Not one player. Not one expected-goals figure, not one box-entry metric, not one line about transfers.
I sat still for a few minutes. Nearly fifty years of recording football taught me something that sounds simple: wrong data is no longer data, it is noise dressed as truth. A wrong label is more dangerous than a gap, because a gap is visible to anyone, while a wrong label drifts quietly into the model and stays there, untraced.
That is why I am writing this. Not to tell the story of a single technical error, but to talk about something those working with football data in Vietnam and Asia often underestimate: the discipline of labeling.
A label is a foundation, not a formality
In the summer of 2026, I learned to trust something no one had yet named: xG. I was 57 that year, working as a transfer-market administrator in Marseille. When Opta first released expected-goals tables for Ligue 1, I did not rush to believe them. I hand-recorded 1,204 shots from 20 teams in the first half of the 2026-18 season and compared them with the actual goals. The correlation coefficient reached 0.84. That number was enough for me to build my own dataset for valuing strikers. Colleagues said my reaction was slow. I needed to verify before using it.

The lesson from that summer did not lie in xG. It lay in the fact that I had to count every shot by hand, assign every shot to a match, a team, a minute. If a shot had been assigned to the wrong match, that 0.84 would have risen or fallen meaninglessly, and I would have built an entire valuation system on sand.
People often think the hardest part of football analytics is the model. I find the hardest part is labeling. A model only does its job correctly with the input it is given. Labeling is the only stage in the whole pipeline that forces us to answer a very human question: where does this item truly belong.
For an item labeled football, I always run three checks. Is there a club. Is there a player. Is there any measurable fact about a match, finances, or competition rules. The film item failed all three. It failed at the very first question.

Anatomy of a wrong label
What caught my attention was not that a wrong item existed, but how it got in. Automated classification systems today read text by semantic pattern: they look at the entities appearing in a piece and see which topic group they sit closest to. That item mentioned a game brand, a major film studio, a director, an actor. In the machine's semantic space, words about gaming, about competition, about performance can sit very close together.
An article about a post-credits scene can touch signals the system associates with 'competitive entertainment', 'esports', 'events'. Once a few matching signals cross a threshold, the label is assigned. The machine does not understand that a Japanese game brand owning a separate film line is entirely different from an esports tournament taking place. To the machine, both are 'content about a competitive entertainment brand'.
This confusion has a technical name: out-of-domain content. The item sits outside the football domain but is labeled football. The problem is not that it is worthless, but that it is worthless to the very domain it claims to belong to.
I once saw something similar on a much smaller scale. In a transfer dataset I received, there was a row showing a young player moving from a second-tier club to a top-tier club for a fee of zero. That row was correct character by character but wrong in essence: it was a player out of contract, not a free transfer. Had I folded that row into my sample for average transfer value, I would have skewed an entire segment. One row, one label.

Three hypotheses before concluding
By nature I dislike concluding early. Before any finding, I always write down at least three hypotheses to explain it, then look for evidence to eliminate each. For that mislabeled item, my three hypotheses were these.
The first: this is a single error, one stray item in a large batch. If so, the share of football-labeled items lacking football entities in later batches would fall to nearly zero. This is the most comfortable hypothesis, and also the one I distrust most.
The second: this is a systematic fault in the classifier, which regularly lumps entertainment content with competitive elements into the football label. If so, the defect lies in the classification, not the operator. The way to test it is to track the error rate across consecutive batches. My self-imposed threshold: if more than 1% of a batch of football-labeled items contains no football entity, it is a systemic problem.
The third: the error comes from a human stage, where someone exported data and let a row from another table slip in. This hypothesis is often ignored because it is less glamorous, but in my trade most data errors are human errors, not machine errors.
What is worrying is that all three hypotheses lead to the same consequence if I do not handle it: that junk item travels downstream.
Downstream contamination: the cost of one bad row
Empty stands are the finest laboratory for a data obsessive. In 2026, when football restarted after the pandemic, I sat in Marseille and analyzed 81 matches played in empty stadiums in the 2026-20 season. Home teams won only 26% of matches, against 43% before the pandemic. I wrote a report on the loss of home advantage. A Ligue 2 club, Le Havre, used that report to negotiate down the price of a young striker who had shone at home.
That story showed me how far data travels downstream. I type a table, a club lowers a contract price, a young player is valued below his true level. Now imagine that same dataset contains one mislabeled row. It does not sit still. It flows into the home-form index, it pumps into the league-wide average, it quietly tilts a buying decision.
In statistics, this is called an outlier. One wrong data point, if extreme enough, can skew an entire regression. The model does not know that point is junk. It only knows that point is an observation, and it will try to explain it at all costs, even sacrificing correct observations.
With a film item slipping into the football label, the worst-case scenario is this. If our pipeline has no domain check, that item moves to deep analysis. A careless person will try to find football meaning in it. They will write about 'competitiveness', 'squad structure', the 'brand strategy' of something that is in fact a film. And so a false signal is born, stored, then cited in another piece.
That is what I fear most in data work: not scarcity, but contamination. With scarcity, we know we are short. With contamination, we think we are rich.
The duty not to fabricate signal
I am 66, old enough to know a number never tells a story unless we ask. And old enough to know there are times when the most honest answer is: cannot assess.
When an item has no football entity, no metric, no match context, then any football analysis of it is fabrication. A decent analyst will not fabricate. They will mark the item as out-of-domain, return it to its proper place, and log the incident for the pipeline operator to fix.
In my trade there is a great temptation that I have seen in many young people: the temptation to prove usefulness by always finding meaning. People are paid to reach conclusions, so they manufacture conclusions. But a conclusion from wrong data is worse than an empty conclusion. An empty conclusion harms nothing. A wrong one does.
There are matches won on the pitch but lost on the data sheet, and I choose the data sheet. But choosing the data sheet does not mean clinging to every number on it. It means trusting only numbers that have been verified, and daring to say 'I do not know' to the rest.
The danger lies in the plausible error
Here I want to argue against a very common reflex in the football-analytics community of Vietnam and Asia: valuing data quantity over label quality.
Many believe that if you collect enough data, the model will find the truth on its own. I do not believe it. A wrong item looks very much like a right item: it sits in the same format, the same field, the same column as thousands of others. It does not incriminate itself. Conversely, an obviously wrong item — a blank row, a missing value — is easy to spot and easy to drop. The real danger lies in the impostor, not the absentee.
This is where correlation does not mean causation within labeling itself. The fact that an item contains keywords close to football correlates with the chance it belongs to football, but that correlation is not strong enough to conclude. Data workers must distinguish between 'looks like' and 'belongs to'. Machines cannot. People must.
I, too, have been the suspect. In 2026, when I proposed using xG to value strikers, many thought I was chasing a fad. I did not argue. I simply gave my sample size and confidence interval. The difference between me and my critics was not belief, but process. They had no verification process. I did.
And here is what I want to say to those building football models: a good model is not the one that reads the most data, but the one that dares to reject data that does not belong to it. The ability to refuse is what separates an analyst from a quoter.
The domain gate: something that should exist
From that incident, I drew a principle and built it into my own process. Before a dataset enters analysis, it must pass a domain gate. That gate asks three questions: is there a football entity, is there a measurable fact, does the label match the assigned domain.
Items that fail the gate are not deleted but redirected to their proper domain. A piece about a film should go into the film pipeline. It has no fault. The pipeline is the one at fault, for accepting it wrongly.
Urban readers often look at the glamour of advanced stats and forget that beneath them lies a layer of raw data. That raw layer is not glamorous. It consists of very boring choices: which team this item belongs to, which minute it is assigned to, which phase it counts for. People measure a midfield's strength through passes allowed per defensive action, and I measure a dataset's strength through its ratio of correct labels.
Croatia winning a tournament with a low PPDA? Then PPDA is just one letter. I wrote that to remind that a single metric is never the truth. The same holds for labels: a single correct label cannot save a contaminated dataset, but a single wrong label can ruin an analysis. The power of the wrong does not lie in its quantity, but in its position. One junk row in the right place can tilt a conclusion about an entire league.
In France, where I work, club data culture is getting tighter. Scouting departments no longer buy a player just because of a highlight video. They buy because of a filtered metric set. Yet even there, I still see transfer datasets with careless labels. No one checks because everyone believes the label is right. Faith in the label is the industry's shared blind spot.
And elsewhere, the esports transfer market suffers the very same disease. Esports organizations label players with unverified metrics, then discount young people based on those metrics. I once saw a player undervalued simply because a dataset mislabeled his role. A click on an esports screen carries the shape of a pass, and it too needs to be labeled correctly.
What I keep
I am not writing this to criticize a specific system. I am writing because that incident is a reminder. As data grows and models grow stronger, discipline at the lowest layer matters more, not less. The more automated the pipeline, the more it must be checked by hand.
A dataset is like a match. We can win on the pitch through one beautiful moment, but if the data sheet after the match is full of junk rows, winning on the pitch is just a blurred memory. I do not want my football dataset to be a blurred memory.
Tomorrow, I will reopen that batch and check every row. Perhaps there was only one wrong item. Perhaps more. Either way, I will write this question into my notebook: if a file about a film can live peacefully inside a football dataset without anyone noticing, how many other junk rows are quietly sitting inside the models we trust?
