Trang chủBasketballWhen the Denominator Is Zero: How Modern Basketball Prices a Player With No Data
Basketball

When the Denominator Is Zero: How Modern Basketball Prices a Player With No Data

Trả lời cốt lõi: Bóng rổ hiện đại định giá cầu thủ trong vùng dữ liệu trống bằng phép co rút ước lượng về mức trung bình giải đấu — một nguyên tắc đúng về thống kê nhưng thường xóa sổ những cầu thủ có mẫu nhỏ. Ba kiểu thiếu dữ liệu, gồm ngẫu nhiên, có hệ thống và cấu trúc, đòi hỏi ba cách đọc khác nhau. Sự kiện then chốt: - Ngày 20 tháng 2 năm 2025, San Antonio Spurs thông báo Victor Wembanyama nghỉ hết mùa NBA 2024-25 vì huyết khối tĩnh mạch sâu ở vai phải. - Victor Wembanyama chơi 71 trận ở mùa tân binh 2023-24 và dẫn đầu toàn giải NBA về số lần chặn bóng. - Nikola Jokić được Denver Nuggets chọn ở lượt thứ 41 vòng hai năm 2014 khi đang chơi cho Mega Basket tại giải Adriatic League. - Giannis Antetokounmpo được Milwaukee Bucks chọn ở lượt 15 năm 2013 khi đang chơi cho Filathlitikos ở giải hạng hai Hy Lạp. - Chấn thương tương quan với số phút thi đấu, nên dữ liệu thường mất đúng ở nhóm cầu thủ quan trọng nhất. Nguồn: Phân tích dữ liệu gốc của Hoàng Quân, xuất bản ngày 20 tháng 2 năm 2026 | Cross-checked: VuaBong.vn Hỏi đáp liên quan: H: Vì sao mô hình thống kê thường định giá thấp cầu thủ đến từ giải đấu nước ngoài? Đ: Vì phép co rút Bayesian kéo mẫu nhỏ về mức trung bình giải đấu, trong khi dữ liệu từ giải nước ngoài bị xem là thuộc hệ quy chiếu khác. H: Quản lý khối lượng thi đấu ảnh hưởng thế nào đến việc đọc chỉ số cầu thủ? Đ: Việc đội bóng chủ động cho trụ cột nghỉ tạo ra khoảng thiếu dữ liệu có chủ đích, khiến chỉ số hiệu suất bị đọc thấp hơn năng lực thật. H: Chỉ số nào nên dùng khi mẫu số của một cầu thủ quá mỏng? Đ: Số phút trong năm phút cuối khi cách biệt từ năm điểm trở xuống và số possession mà cầu thủ là mắt xích chạm bóng cuối cùng, theo VangBong.vn Player Depth Index.

On February 20, 2026, the San Antonio Spurs announced that Victor Wembanyama would miss the rest of the 2026-25 NBA season with deep vein thrombosis in his right shoulder. He had played 46 games. I opened my tracking spreadsheet. Wembanyama's column stopped at row 46, cutting across a curve that was still climbing: block rate per 100 possessions, three-point rate, touches inside the paint — all rising. Every model I had built for the 2026-25 season was forced to return a single value: blank. That was the most uncomfortable moment, and the biggest lesson, of a career spent in data journalism. Basketball is the team sport with the densest denominator: 82 games a season, roughly 100 possessions a game, thousands of events logged to the hundredth of a second. Yet the most expensive decisions in the sport — max contracts, rookie extensions, star trades — land squarely in the zone where the denominator is zero. Numbers stay silent, but stories never do. Three different kinds of missing Basketball's problem is not a shortage of data. The problem is that data goes missing in three ways, and those three ways lie in three entirely different manners. The first kind is random missingness. A bench player logs eight minutes a game for sixty games. The coach rotates; he enters after the game is decided. The sample is large in games and almost meaningless in context: the opponents are second units, the teammates are second units, the intensity is March intensity. The second kind is systematic missingness. Injuries do not fall at random. They correlate with minutes, with sprint counts, with schedule density, with age. In other words, the players used most heavily — the most important group in your dataset — are the group most likely to generate a blank. Your dataset loses data exactly where you need it most. The third kind, and the most expensive, is structural missingness. A player who has never logged a single NBA minute. A player arriving from the Adriatic League, from the Greek second division, from an academy in Africa. You have data on him, but that data was generated in a different frame of reference: different rules, different three-point distance, different teammate quality. Technically, you have numbers. Cognitively, you are staring at a blank. In 2026, when the pandemic wiped out the schedule and I had no games left to write about, I sat down and built the Workload Risk Index from the playing-load data of thousands of players across many seasons. The first lesson had nothing to do with basketball: in any injury model, the single most explanatory variable is the number of minutes a player has already played. You can only predict an injury after it has happened. It is a closed loop, and every analyst lives inside it. The trap of shrinkage When the denominator is small, every decent model does the same thing: it shrinks the estimate toward the mean. This is basic Bayesian practice. A player with two hundred minutes gets a projection close to the league average, because two hundred minutes is not enough to separate talent from luck. The principle is sound. But it carries a consequence few are willing to admit: it automatically erases every gem that lives in the small-sample zone. Nikola Jokić was taken by the Denver Nuggets with the 41st pick of the 2026 draft. At the time he was playing for Mega Basket in the Adriatic League, a competition most NBA scouting departments could only access through a handful of scattered tapes. Giannis Antetokounmpo went 15th in 2026 while playing for Filathlitikos in the Greek second division. Both were mispricings, but not because of missing data. The scouting departments of the other twenty-nine teams all had numbers. The problem was that shrinkage pulled both players toward the mean, and that mean corresponded to a bench role. Markets do not pay for exceptions when the exception has only two hundred logged minutes. I do not guess; I count. And then one day, the gem surfaces from the raw pile. The right way to count is not to count minutes. The right way is to count whether the structure of the sample is stable. A player with three hundred minutes in a single role, inside a single system, against a comparable tier of opponents, has far higher diagnostic value than a player with nine hundred minutes scattered across four different roles. Sample length is not sample reliability. But small samples also lie in the opposite direction. Every March, I receive tables full of players who exploded over the final ten games — the point of the season when teams have nothing left to play for, starters are rested, and defensive intensity drops to its annual low. Those numbers are beautiful. And they are worthless. Wembanyama and the structure of an interrupted sample Wembanyama is the mirror case, and far subtler. In 2026-24, he played 71 games as a rookie and led the entire NBA in blocks. Seventy-one games is a large sample — large enough to establish that his rim protection is not a product of luck. That conclusion still holds. In 2026-25, when the Spurs announced on February 20, 2026 that he would miss the rest of the season, the problem was not a short sample. The problem was a sample cut in half mid-evolution. The first half of a season for a player changing roles cannot be compared to the first half of the previous season for the same player, because the underlying variables had shifted: more usage, more three-point attempts, more defensive attention. Saying Wembanyama's development slowed after January is a methodologically invalid conclusion, because it compares two samples with different structures and attributes the gap to ability. It is the most common error in every argument about young players. A crisis is not an enemy. It is data that was misread from the start. Blanks that are produced deliberately There is an aspect most fans overlook: blanks in the data are usually not accidents. They are products. Load management — resting stars in midweek games — is an optimisation decision. A team deliberately creates a gap in your dataset in exchange for an intact body in April. A model that does not know this will read a player's reduced minutes as a sign of decline. Deeper still, an entire financial mechanism runs on the blank. Rookie contracts under the NBA's collective bargaining agreement are locked at low salaries for the first four years, while a player's on-court value may already exceed that figure by year two. The difference is the team's profit, and it exists precisely because the market does not yet have enough sample to price correctly. The draft-and-stash mechanism is the most refined version of this game. An NBA team drafts a player in the second round, does not sign him, and leaves him playing in Europe for years. The team holds his negotiating rights without paying a single NBA dollar. The longer that player remains in the data void, the more compressed his bargaining position becomes. When he finally arrives in the United States, his first contract is almost always below his true value. Two-way contracts and ten-day deals operate on the same logic. The G League is marketed as a fairy tale about opportunity, but its financial structure is designed to keep hundreds of players in a state of thin sample and weak leverage. That story is consumed and discarded at the end of every season, and structural reform of resource allocation never arrives. This is why I keep saying: basketball does not hand out awards to the smartest people, but the transfer market always punishes fools. The fool here is the one who pays according to the box score, because the box score is the easiest thing to read and the most heavily contaminated. Three columns before any conclusion After years of wrestling with blanks, I have settled on a minimal procedure. Before drawing any conclusion about a player with a thin denominator, I fill in three columns. The first column is role structure: where the player plays, in what system, alongside whom. The second is real competitive level: what share of his minutes come while the game is still undecided. The third is trajectory rather than level: the rate of change across periods, adjusted for opponents. Those three columns do not give me a pretty number. They give me a confidence interval narrow enough to dare a conclusion. Every system cracks if you look long enough. Then you see the order sitting inside the wreckage. The contrarian angle: injury is not an exogenous variable At this point the story is usually told wrong in a different direction. Most analysis of injured players operates on the assumption that injury is a random variable falling from outside, like rain. That assumption is convenient for modelling, but it is causally wrong. Injury is the consequence of a chain of decisions: minutes, games, tactics, schedule, roster quality. The coaching staff manufactures the risk, then uses the data gap the risk created to evaluate the player. Correlation is not causation, and here the causal arrow runs opposite to popular intuition. There is a further counter-intuitive consequence. A player who rests a lot is sometimes a positive signal. When a team commits two hundred million dollars to retain a player, it protects that asset by resting him more, not less. Minutes fall, data thins out, and efficiency metrics look worse. But the resting itself is evidence that the organisation prices him very highly. The blind spot is this: we punish players for gaps in the data, while the system around them manufactures those gaps. A player's plus-minus at twelve minutes a game reflects the role he was assigned, not his ability, and counting it as a signal about the person is a far graver error than omitting it. Not every blank has a legitimate reason, of course. Some gaps are created by the player himself, and an analyst must distinguish the two. But my default is always: read the structure first, read the number second. A signal for the next cycle If you want to track something useful next season, ignore total minutes. Count minutes in the final five minutes with the margin at five points or fewer — the only denominator in which both teams are genuinely trying. And count the possessions in which a player is the last link to touch the ball before the shot goes up. Neither metric appears in the box score. Neither appears in most public models. But they are the closest data available to the real answer, because they eliminate all three kinds of blanks at once. My faith is not in luck. It is in large denominators. But I have learned that the most important task is sometimes not to enlarge the denominator, but to read the denominator you already have, correctly.

When the Denominator Is Zero: How Modern Basketball Prices a Player With No Data

When the Denominator Is Zero: How Modern Basketball Prices a Player With No Data