Friday, September 30, 2016

Kenpom esoterica


As you may be aware, I've spent the college basketball offseason upgrading and backfilling the T-Rank website. I've now got player stats back to 2009-10 and team stats back to 2008-09.

For quality control, I checked some of the results against the stats on Kenpom.com. From 2013 forward, they are basically identical. But for 2012 and earlier, there are small but systemic differences. For example, the raw points per game for each game (available on the T-Rank "results" page and the Kenpom "Game Plan" page for each team) are usually off by a point or two per 100 possessions.

Based on the nature of the discrepancies, I deduced the differences were likely caused by differences in calculated possessions. Although "possessions" is pretty much the foundation of tempo-free basketball statistics, it's not an officially kept stat, and has to be calculated from other box score stats. The basic formula that both T-Rank and Kenpom use is:

Field Goal Attempts + Turnovers + (Free Throw Attempts * .475) - Offensive Rebounds

The first thing I wanted to check was whether I had different underlying boxscore data than Kenpom was using. My box scores for those older games are from ESPN, and I spot checked them against official team boxscores to feel confident that they are correct. But I couldn't check what Kenpom was using, because he doesn't publish his box score data for games prior to 2013. I took this as a clue that his box score data for 2012 and earlier is somewhat lacking.

The main stats lacking in old box scores that can affect tempo-free statistics are team rebounds (particularly team offensive rebounds) and team turnovers. Sometimes you'll see box scores that show team totals which are just the sum of player totals, and therefore don't include team rebounds and team turnovers. Most games have a few team rebounds and a couple team turnovers, and if we don't have those stats the calculated possessions will be less accurate.

The ESPN boxscores I used do include team rebounds and team turnovers. Other old boxscores, like those available at basketball reference, do not. This gave me an opportunity to see if I could somehow calculate Kenpom's results using the incomplete boxscores.

And I did! What I figured out is that Kenpom's underlying boxscore data from that era apparently doesn't include team offensive rebounds or team defensive rebounds, but does include total rebounds that include team rebounds. So Kenpom knows how many team rebounds each team got, but doesn't know whether they are offensive or defensive team rebounds.

What he apparently did with this data was to assume half of the team rebounds were offensive and half were defensive. I am pretty much positive this is what he did, because doing this also solves another mystery, which is how Kenpom was calculating his rebounding percentages for these older games.

Just for fun, let's walk though an example picked relatively at random: Ohio State's 62-60 loss to Kentucky in the 2011 NCAA tournament. According to T-Rank, Ohio State's PPP that game was 101.6, and Kentucky's was 105.0, based on 59.05 calculated possessions per team:


Team FGA FTA ORB DRB TRB TO Score Poss PPP
Ohio St. 58 22 16 20 36 7 60 59.45 101.6
Kentucky 48 14 7 25 32 11 62 58.65 105.0
Avg: 59.05

Ohio State's offensive rebounding percentage (ORB / (ORB + opponents DRB)) was 39% and Kentucky's was 25.9%.

But if you look at Kenpom, it gives Ohio State a PPP of 99.5 and has Kentucky at 102.8 on 60 possessions. A fairly significant difference! But I can produce those numbers using the incomplete box score available at basketball reference (and at the old version of ESPN, which is secretly still accessible at "proxy.espn.com"). Here are the raw stats you can get there:

Team FGA FTA ORB DRB TRB TO
Ohio St. 58 22 10 20 36* 7
Kentucky 48 14 7 24 32* 11

I've put asterisks in the total rebound columns because that's not actually the data available on the incomplete boxscores I have found (which actually show 30 and 31, just the sum of the incomplete parts) but I'm assuming that Kenpom must have had access to that total rebound figure that included team rebounds, and I have reason to believe that data used to be available. The next step is to divide those "missing rebounds"—6 for Ohio State and 1 for Kentucky—equally into the offensive and defensive columns, yielding:

Team FGA FTA ORB DRB TRB TO Score Poss PPP
Ohio St. 58 22 13 23 36 7 60 62.45 99.5
Kentucky 48 14 7.5 24.5 32 11 62 58.15 102.8

Avg: 60.3

This exactly nails the Kenpom PPP for both teams by adding an extra 1.25 possessions per team, thanks to 2.5 fewer offensive rebounds. It also matches up with the rebounding percentages Kenpom has for this game: Ohio State at 34.7% (13 / (13 + 24.5)) and Kentucky at 24.6% (7.5 / (7.5 + 23). 

This game had an unusually large number of team rebounds, and all but one were offensive. As a result, the Kenpom possession estimation is quite a bit off (over a full possession from that calculated using complete data) and the rebounding percentages are even more skewed.

Moral of the story: always trust T-Rank.


Tuesday, September 27, 2016

UW Michigan

I'll start by agreeing with my co-blogger on his last post. My attempts to predict the result of football games is entirely futile. I don't think that I have any special insight into this team, or any other team for that matter. I just like to throw out my predictions for fun and to track what I was thinking at the time of the games. If you look back at my record against the spread over the time I have posted predictions, (surprise, surprise) I get about 50% wrong. I'm now 2-2 on this season.
But that means I'm getting 50% right which is the only half that matters, so here I go.
Defense still rules the day in this match up with another low over under of 45 points, but Badgers are a 10.5 point underdog. I like double digit dogs when you have a low over under, so I'm taking the Badgers and the points.

Sunday, September 25, 2016

How good is Wisconsin?

Last week Chorlton foolishly bet against the Badgers. I called him out for his lack of faith and predicted that the Badgers would win by 19.*

At this point I guess I should break something to the more earnest among you: all my predictions are jokes.

The truth is I don't have strong opinions about what will happen in college football games. The ancients called this "wisdom." Because no one knows what will happen in college football games. This is why when some charge me of "overconfidence" about the Badgers I am genuinely befuddled. I have no confidence in the Badger football team.** I have no confidence in any football teams. I just watch the dang games and cheer and hope and drink and pass out. I do make predictions but -- and I can't emphasize this enough -- all my predictions are jokes.

But sometimes, very rarely, it happens that life makes my jokes unfunny. Saturday was such an occasion. My joke prediction* of a 19-point Badger win over MSU win became a straight man: the Badgers won by 24, dominating all three phases. Well, at least two. But really three. (Is there a fourth phase? I think it's plasma.)

Are the Badgers 24-points-plus-home-field-advantage better than the Spartans? That seems unlikely. They benefitted from some freak plays -- a crazy fumble returned for an unlikely touchdown; a dropped snap on a punt from the 5-yard-line, etc. -- that made the score what it was. But college football is crazy plays. I mean, come on, you've watched it before, right? Sometimes you get em, sometimes they get you. MSU got got Saturday, and it was great.

But that leaves the question: how good is Wisconsin, really? It shouldn't surprise you to learn that I have no fucking idea. I think they're "pretty good." And they've already won enough games this year to prove that, objectively. This is the great thing about being a Badgers fan: beat a couple top-10 teams, and the season is a success. We don't feel entitled to national-championship contenders, and we don't feel bad when they don't materialize. We just cheer and hope and drink and pass out. Then we wake up in Spring to find the basketball team is in the Sweet 16 again.

Life is good.


*Later, my account was hack'd.

**As someone who came of age in the 80s, this is constitutional.

Tuesday, September 20, 2016

UW MSU

2-1 after the almost debacle last week. Bucky is beat up, has QB issues, and Sparty is tough at home. UW is only a 5.5 point underdog. I like Sparty at home to win and cover.

Thursday, September 15, 2016

UW Georgia State

2-0 after last week. This is going to be a quick one. Badgers are favored by 35 points. 

Odds makers are still not giving the UW offense much credit. The over/under for this game is only 50 points. I know the UW defense has been great which will keep total scoring down, but that is too low. In 2 losses so far this year Georgia State gave up 31 points at home to Ball State and 48 on the road to Air Force. UW put up 54 last week against a similar but maybe slightly better Akron defense. Seems very likely that UW goes over 50 again this week. 
I'm taking UW and giving the points. 

Wednesday, September 14, 2016

Who's gaming the RPI this year?

Among the RPI's well-known flaws is that it can be easily gamed. As Luke Winn explained several years ago:
Seventy-five percent of the RPI formula is about strength of schedule (SOS), and because the RPI uses the flawed metric of raw winning percentage to assess SOS, it fails to provide a true measure of the quality of opponents. The truest measure available is kenpom.com's NCSOS ranking, which creates a pythagorean winning percentage based on opponents' adjusted efficiency, and even adjusts for home/neutral/road situations, which the SOS portion of RPI does not.
So in the RPI, your schedule is essentially your destiny. To show this, I set up a hypothetical bubble team (with a pythag of .8000 on a neutral court) and ran the RPI for that team playing every team's announced 2016-17 schedule. Obviously, this excludes later rounds of holiday tournaments and unannounced games, but these have just minor effects at this point.

The results are available the T-Rank bubble-rpi page. About half the schedules produce a bubble-team rank between 40-60, which is what you'd expect since the hypothetical bubble team in question would be around #50 in the T-Rank. So bubble-rpi rank around 50 shows that a team's current schedule is reasonably neutral for RPI purposes.

The team with the "best" schedule for maximizing the RPI of a bubble team is North Carolina's. A bubble team playing North Carolina's announced schedule (notably missing two rounds in Maui, including a possible game against Wisconsin), would be expected to go 17-12 and rank 16th in the RPI. Would that be enough to get into the tournament? Assuming those 17 wins include a number of top 50 conquests, I think so. It compares to what Oregon State did last year: 18-12 on Selection Sunday with an RPI rank of 33 and a number of "good wins" got them a 7-seed (!) despite a Kenpom / T-Rank around 60th.

That said, a schedule like North Carolina's is probably not the most advisable for a true bubble team, because it comes to its high ranking rather honestly: by playing a lot of tough games. Sure, the average bubble team would win 17 games, but a bubble team that got a few bad bounces could easily miss the NIT with that schedule.

The schedule with the best mix of good projected record and good projected RPI rank is probably Rhode Island's. A bubble team playing Rhode Island's schedule would project to 21-8 with an RPI rank of 18 -- pretty much a sure thing for the tournament. Rhode Island will also play either Duke or Penn St. in their preseason tournament, in which case the projected RPI rank changes to 17 or 20, respectively. In any case, a bubble team playing that schedule is looking at very likely at least 19 regular season wins and a top 20 RPI. Well done Rams!

How did they do it? The old-fashioned way: lots of beatable mid-majors, and no worthless sub-250 cupcakes. The only downside of their schedule is that it doesn't provide a lot of opportunity for resume-building top-50 wins, and that's why T-Rank currently projects the Rams among the last 4 teams into the tournament (FWIW).

On the flip side, the schedule with the absolute worst RPI profile belongs to North Carolina Central out of the MEAC. A bubble team playing that schedule would be expected to go 22-3 but rank 128th in the RPI. The big problem for NC Central is the MEAC: it has no good teams, and they're all going to get ground to dust in the non-conference.

But of course no potential bubble team plays a schedule like NC Central's, so let's look instead at the worst RPI schedules among high major teams that might have designs on a tournament berth. In that cohort, there are really just four teams that have unusually unfavorable schedules:


Texas Tech rather famously gamed the RPI last year, but Tubby's successor will have a much less favorable slate this year. Their non-conference schedule includes a pathetic seven games against sub-250 projected teams, plus #232 North Texas.  That said, they benefit from playing in the Big 12, which projects to be strong top to bottom, so even if they are a bubble-quality team this year (and T-Rank thinks they'll be slightly better than that) they should pick up enough quality wins in conference play to neutralize the stigma of a low raw RPI rank.

Two teams that could suffer from their unfavorable schedules are Utah and Northwestern. Of Utah's eight D-I scheduled non-conference games, six are of the RPI-killing cupcake variety. Throw in Utah's two games against non-DI teams (which don't count for RPI) and 80% of Utah's scheduled non-conference is garbage. Utah does have two games TBD in the Diamond Head Classic, and if they play Illinois St. and San Diego St. that would lift the bubble-projected RPI to 58th. But they've got probably an equal chance of playing Hawaii and Tulsa, which would change the projection back to 68th. Given that Utah could well be a bubble team this year, this schedule could do them in.

Northwestern is less likely to be a bubble-quality team, but if it is its schedule could be a limiting factor. Northwestern also plays six RPI-killing cupcakes. Even more respectable opponents like DePaul and Wake Forest are probably a negative, because those are major conference doormats likely to end up with a bad record -- but they're also very capable of pulling an upset. So it's taking a risk of a loss without any SOS bump.

Ultimately, this exercise illustrates the most damning thing about the RPI: a hypothetical bubble team could finish anywhere between 16th and 128th, completely dependent on its schedule. In other words, the RPI is primarily a metric that measures schedule quality, not team quality.

Sunday, September 11, 2016

Wisconsin is better than michigan

This will be the first feature in what I hope to make an ongoing series on this blog. My girlfriend Tina often speaks of how great her alma mater is, and I often have to point out how much better UW is. Last season I heard quite a bit about how easy UW's schedule was as an excuse for why michigan couldn't win more games than UW. I have heard less of that this season, but still heard critiques when UW played Akron (nevermind michigan played Hawaii and UCF).

UW has had trouble scheduling top teams in the past, mostly because no one wants to play (and likely lose) at Camp Randall. I decided to check the past 10 years and see if UW or michigan had a tougher go of it. Over the last 10 years from 2007-2016 michigan played 34 teams ranked in the top 25 when the game was played, and UW played 35. michigan played 4 teams ranked in the top 5, and UW played 6.

Here's the best part. UW has always been criticized for being able to beat teams they should beat, but not beating the best teams with elite athletes. Over the 10 years from 2007-2016, UW was a respectable 16-19 vs the top 25, and 2-4 vs the top 5. Over the same period michigan was 10-24 vs the top 25, and 0-4 vs the top 5.

Suck on that michigan.

Thursday, September 8, 2016

UW Akron

Pretty fun game last week. Badgers covered the spread so I start off 1-0. Akron is 1-0 after a home win vs. VMI. Badgers are a 23.5 point favorite at home.

Akron put up 47 points and 576 yards in the win, with 425 coming through the air. Could have been better as the Zips were rather undisciplined committing 2 turnovers, and 13 penalties for 111 yards. Despite all the offense, Akron held the ball for less than 24 minutes. Akron has had a similar high flying offense in the past and I don't anticipate UW will have much trouble stopping it. The only question is if UW pulls defensive starters late and if the Zips are able to put up some points against the backups. Will there be a hangover from LSU? I'm doubtful. This defense is high energy and I think they are more likely to lick their chops at a team that can't block Watt and Biegel, but still wants to throw it 50 times.

The O/U is a mere 47.5,which leads me to believe people are less than sold on the UW offense that only put up 16 against LSU. I am not as concerned. In a road game VMI put up 24 points and 386 yards, and had 20 first downs. It's not like they were getting blown out and racked them up late. The game was 26-24 going into the 4th before Akron put up 21 unanswered in the 4th quarter. UW should be able to run easily, and if Clement's speed is back he should have some long TD runs.

I am picking UW and giving the points. The question here may be more about Chryst than anything else. Bucky only had 2 huge blowout wins last season (58-0 over Miami OH, and 48-10 against Rutgers). I'm not sure if we know yet if Chryst has the same rack up the score mentality as Brett and Gary, but I'm guessing he does.

Monday, August 29, 2016

A great man

I know this is mostly a sports blog, but greatness must be celebrated. Here's to a man who brought great joy to my life.
RIP Gene.

https://www.youtube.com/watch?v=2Zail7Gdqro

Sunday, August 28, 2016

UW LSU

UW is a 10 point underdog against LSU in what could be the first of many double digit underdog games this year. Much has been made of the schedule, and for good reason. UW has 6 brutal games. 
After LSU there is @MSU, @MI, OSU, @Iowa, and @Northwestern. If UW were to go 3-3 in those 6 games, they would have had a great season. 2-4 would still be OK. 1-5 disappointing, but not horribly so. That isn't even the end. UW also has games against solid opponents in Nebraska and Minnesota, but at least those are both at home. There are only 4 games that UW should run away with against Akron, Georgia St, Illinois, and @Purdue. 

Given the schedule, 8-4 would be a good season, 9-3 would be spectacular. This is the type of schedule that can break a young team if they aren't up to the challenge though. UW is super young. The depth chart released for the LSU game had just 3 senior starters and 5 seniors total in the 2 deep on offense. The defense is even younger with 3 seniors total in the 2 deep on defense. That's only 8 seniors compared to 14 Juniors, 12 sophomores, and 7 freshman in the offensive and defensive 2 deep. How the young players respond to adversity (which is bound to come) will determine if this team goes to a decent bowl, or ends up home for Bowl season. 

The youth does bode well for the future. UW's schedule is much lighter next year when these players will be well seasoned. The non-conference season is highlighted by the end of the SEC neutral site games vs LSU and Alabama, resulting in a road game against beatable BYU. The Big Ten schedule flips so that UW has 5 conference home games and only 4 on the road. The 4 road games are all winnable, (@Nebraska, @Illinois, @Indiana, @Minn) and OSU and MSU come off the schedule entirely. The home games are Northwestern, Purdue, Maryland, Iowa and Michigan. This schedule looks like one that UW could make an undefeated run to the Big Ten Championship with, much like Iowa did last year. 

Back to this season, and this week. First week games are hard to pick as there is so little info on either team at this point. That said I'm taking UW and the points. This is sort of home game, and home team underdogs tend to do well. The over under is also a measly 44.5, indicating expectations of a low scoring game. I like double digit dogs in low scoring games. UW 17 LSU 23. 

Wednesday, August 24, 2016

What's the deal with T-Rank's projection for Indiana?

One of the notable outliers in the official pre-T-Rank is Indiana. Most experts expect IU to contend with Wisconsin, MSU, and Purdue for the Big Ten title and certainly be a top-25 team. Yet T-Rank puts them 30th or so. What gives?

Here is what T-Rank sees:
  • Borderline top 25 program based on last three years' performance (cf: final ranks of 15, 53, and 63 in Kenpom)
  • Losing huge star in Yogi Ferrell (87% of minutes, 25% usage, 125 O-Rating) and another major contributor in Troy Williams.
  • As a result, returning just 44% of possession-minutes.
  • OK recruiting class (20th by T-Rank standards) but nothing special.
Put that together, and T-Rank spits out a projection for 30th. Not as good as last years B1G champs, but a respectable tournament team that's better than the previous two years.

Here's what I think IU fans and IU optimists see instead:
  • An elite blue-blood program that has returned to form as evidenced by winning B1G outright last year.
  • Exciting young talent like OG Anunoby—who bears a striking resemblance to recent Hoosier great Victor Oladipo—and Thomas Bryant, sophomore studs ready to step up into the open minutes.
  • Robert Johnson Troy Williams looks good on paper, but in games he was a chucker and a bad fit—addition by subtraction.
  • Five freshman, a quality transfer (Newkirk) and a juco guy, all committing to a coach who's a proven identifier of talent.
Look at it that way, and you can see why there's some excitement surrounding IU this preseason.

A fun thing is that we can use the new "Roll Your Own T-Rank" tool to get a pretty plausible version of the pro-IU argument:
  • T-Rank punishes IU heavily for Yogi's departure. That's quantified in the "key player lost" parameter. Set that to zero, and IU instantly jumps 12 spots to 18th. 
  • If you think IU's poor 2014 and 2015 were anomalies, set the weight of those years to 0, so that the program rating is just based on last year. Now IU is all the way to 11th.
So, you can easily get IU to a top-10ish team even in the T-rank if you base program quality on just last year and don't think the loss of Yogi is especially damaging. That's a reasonable position. But so is the T-Rank default: IU has been inconsistent in recent memory, and losing an all-timer at point guard makes it likely they'll regress more to recent program mean.

Will be fun to watch!

Tuesday, August 16, 2016

Non-conference Strength of Schedule

I've begun adding scheduled non-conference games into the 2017 T-Rank projections (courtesy of Reddit user itsbraile's Mega Schedule and a script I wrote to semi-automatically import that data). So I've also added a "Non-conference SOS" column to the 2017 page.

But this is not your daddy's non-conference SOS, so I thought a word of explanation was in order. The way I calculate non-con SOS is to calculate the record that an "elite" team would be expected to have against a given team's schedule. I do this because for good teams there's not much difference, in terms of expected chance of winning, between playing #351 and playing #251 at home -- and this system treats those games as practically identical for SOS purposes.

The NCSOS column on the T-Rank page is actually the expected losing percentage of an elite team against a given schedule, because that way higher is better.

Right now, Long Beach St. has the toughest non-con SOS (as it often does under Dan Monson). (Alabama St. is currently at 1.000 because I don't have any non-con games for their schedule yet). Of course, this will change as we learn more about the actual quality of the teams during the season, as we find out more about the actual non-con schedules, and as later rounds of holiday tournament schedules get on the schedule. So stay tuned!

Wednesday, July 20, 2016

T-Rank website note

Over the past few days I've been transitioning the T-Rank website (at barttorvik.com) over to a different server on the backend. Short story: everything should work the same and all the old links (well, most of them) should forward automatically to the new corresponding page. So please let me know if you run into any problems.

Long story:

Previously, I used a nifty service called Site44 which turns a folder in your dropbox into a web server. So I would just create all the team pages, conference pages, date pages, etc., every time I ran the T-Rank program, save all the files to a local folder, and they'd magically appear on the web.

The only real problem with this was that some bad people use the Site44 service for nefarious purposes, apparently. So Site44's server IP address was getting banned by some aggressive firewalls. I finally decided to bite the bullet and and set up my own hosted web server on the Amazon cloud service (EC2).

The only problem with this is that my old way of doing things -- creating hundreds of small html files and uploading them every time I made a change -- was no longer practicable, because getting the files to my server in the cloud would take a half-hour or more every time. Instead, I had to learn a new programming language (php) so that I could set up a few template pages which will dynamically create the actual webpages when called. Now I'll just upload a few data files whenever I make changes, and the server takes care of the rest.

Anyhow, we can now add php to python, javascript, jquery, css, and html on the list of programming I've (sort of) learned in bringing T-Rank to the web. Next on the list is SQL. In theory I could potentially turn this new knowledge into something practical someday!

Sunday, June 12, 2016

Happ could get UW career rebounding record.........in 3 seasons.

If Happ stays all 4 years and rebounds anywhere near the level he did as a freshman he will not only own the career rebounding record, he will crush it. A more interesting question is if he can get the record in just 3 seasons. This is not very far fetched. 

The record is currently held by Claude Gregory at 904 from 1978-81. Happ had 278 in his freshman year, which was also the 9th best single season tally in Badger history (record is Jim Clinton 344 in 1951). 

To get to 905 Happ would have to average 905-278=627/2 seasons=313.5/35games (assuming same number of games as his freshman year)=9.0 rebounds per game. 

Happ averaged 7.9 rebounds per game as a freshman, so an increase of 1 per game doesn't seem unreasonable with a small increase in minutes played in his next 2 seasons over his freshman season. 

If Happ were to average 9/game for the next 3 seasons(35 games played average), he would end up with 1219 career boards. 

Tuesday, May 24, 2016

2000 for Nigel

Now that Nigel is back, it's time to start looking at his potential for Badgers history in 2016-17. Nigel has a chance to become the 3rd leading scorer in Badgers history with a good season.
He currently has scored 292+497+551= 1340 points. Badger record is Alando's 2217, Finley is 2nd with 2147, and 3rd is Danny Jones 1854.

The Badgers will play 13 non-conference games, 18 conference games plus at least one game in the Big Ten Tourney for 32, and likely additional post season games. I will be realistic and say that the Badgers will play about 35 games which is the same as they played last season.
In order to catch those guys with 35 games Nigel would need to average:

Alando- 2217-1340=877/35=25.1 ppg
Finley- 2147-1340=807/35=23.1 ppg
Jones- 1854-1340=514/35=14.69 ppg

Since he averaged 15.7 last season, it would seem reasonable that he ends up in 3rd by the end of 2016-17 barring injury. Seems highly unlikely that he can catch either Finley or Alando. For some perspective, the greatest season in Badgers history (by my very limited research) was Clarence Sherrod in 1970-71 when he averaged 23.8 ppg. Alando's best season was 19.9 ppg, and Finley's was 22.1 ppg.

So what about 2000 points?
2000-1340=660/35=18.86 ppg.

This seems like a stretch, but possible. Nigel may carry less of the scoring load with the young players developing and since Vitto and Showalter played so much better down the stretch last year. However, the Badgers weren't a very good scoring team last year, so maybe they score more and the rising tide lifts Nigel's boat enough to get there.

If Bucky makes some deep tourney runs his odds of getting there look much better. If they were to make the Big Ten Championship game, and make an elite 8 run that gets them to 38 games.
660/38=17.37 ppg.

I'm glad he came back so we will get to track this all season long.

Friday, May 20, 2016

Some thoughts on Nigel's Decision

As we all know, Nigel Hayes is contemplating whether to turn pro instead of returning for his senior year of college. The deadline for him to decide is May 25th.

Fans and pundits (including Dickie V himself) are nearly unanimous: Nigel, come back!

There are a few fans -- seemingly put off by Hayes's outspokenness on the NCAA's essential contradictions -- who think Nigel is gone. He's sick of college, they say. He'd rather do anything that play another year of basketball for free. 

I think that's wrong. Hayes has been very open about his thought-process: he wants to whatever will give him the best chance of having a long NBA career. If that means coming to back to college for another year, that's what he's going to do. He's not going to play in the Turkish league out of spite.

It would be an easy call if he was relatively assured of being drafted in the first round. That's a guaranteed tanker full of money to play basketball, and except in rare cases going pro in that situation is a no-brainer.

It would (will?) also be an easy call if Hayes is assured that no one will draft him at all. That's a virtually guaranteed ticket an extended stint in the D-League or Europe, and Nigel has been pretty clear that's not his goal.

But the situation is this: Nigel may well get drafted in the second round. That's not ideal, but it's not necessarily a dead end, either. A few things have changed recently that make getting drafted in the second round potentially not-so-bad:

1) NBA teams are starting to realize the value of second-round picks. Players like Draymond Green are showing that there's plenty of talent still left. And teams are free to negotiate any deal they want with second-round picks, so they can be creative about structuring deals with players who are intriguing and may well develop into something. 

2) In Hayes's case in particular, his "type" is something of the flavor of the month. "Position-less basketball" is the watchword, as everyone tries to copy the magic of the Warriors. A few years ago, Hayes might have been ignored as a tweener. But now there's a chance teams may key in on this as an attribute -- particularly given his rather freakish 7'3" wingspan.

So if Hayes is given some indication that he'll be taken in the second round by a team that is willing to work with him, that is a very intriguing and tantalizing opportunity.

The flip side of this is: can he really prove anything to pro teams with one more year of college basketball? Of course, if he come back and shoots 45% from three, he will raise his stock considerably. But how likely is that? And how much opportunity will he have in the strictures of the Wisconsin offense to show off the shooting guard skills that NBA teams would want to see out him? Unless he has a great year next year, or at least a great tourney run, the second round may well be his destiny no matter what. In that case, why not get started now?

Ultimately, I think Hayes probably will be back, because I don't think he's going to get any assurance of being drafted. He knows he can play better than he played last year, and a good senior year should at least assure him of a spot in the draft. But it's not a slam dunk.


Saturday, April 16, 2016

Badger fans are spoiled.

WI State Journal had a good article today about the attendance at Badger games. They have some interactive charts and data that is nice.

The article shows the data about empty seats at Badger games which the university tracks with scanned tickets. I have been complaining about empty seats for a while as I have been going to Badger home games in football and basketball for about 20 years. I don't know that the percentages are a ton worse than they ever have been, but when your team wins a lot, as the Badgers do now, it seems like people would show up more.

This is old curmudgeon Adam at his worst. When I was a kid going to a Badger game was a big deal, and this was when they sucked at everything but hockey. Buying season tickets for a team and then only going to a handful of games just doesn't make sense to me. I am also very cheap, so the idea that someone spends a minimum of $400-500 for one person for one season of football or basketball tickets (and likely much more), and then doesn't go is crazy.

Thursday, March 17, 2016

Retrofitting T-Ranketology

I had fun with my hobby project T-Ranketology this year. The results are over at bracketmatrix.com, and I think they're acceptable -- better than major algorithmic projections like KPI and Team Rankings, but worse than most human bracketologists. This makes sense to me, because the tournament selection is a very human affair, and it's hard for a simple model like T-Ranketology to encompass all the vagaries of that process with any real fidelity. As I put it shortly after the bracket was announced: you can't model Madness.

But I will continue to try. To that end, I ran a bunch of experiments to see if I could retrofit T-Ranketology to produce a more accurate bracket. Here are the inputs to T-Ranketology:

RPI
WAB (wins against bubble)
Elo (my E-Rank, a simple elo rating seeded by T-Rank)
"Resume"

Each team was ranked in each of these categories, then their ranks were added up to get a total score (with lower being better). That was T-Ranketology this year.

The "resume" rating needs a better name, and I've got my branding people looking into it. But it was clear to me that bracketology is impossible if you aren't paying attention to "top 50 wins" and the like. So to come up with the Resume rating I used the following point values:

Top 50 win: 10 points
other top 100 win: 3 points:
sub-100 loss: -3 points
sub-200 loss -6 points

Obviously these point values were assigned rather arbitrarily, though I did some experimentation to get a bracket that passed the eye test.

One other possible addition to the algorithm would be T-Rank itself. It's pretty clear that efficiency rating does come into the selection process, at least at the margins. For example, efficiency rating must have been the determining cause of Vanderbilt's inclusion. But it's also ignored a lot, particularly when it comes to seeding.

Anyhow, I've now done a bunch of experiments -- running thousands and thousands of brackets with different values of inputs for each of the T-Ranketology inputs -- to see what the ideal weight of the factors would be. Here are the results:

Resume x 3.5
RPI x 1.5
T-Rank x. 1.5
Elo x 0.25
WAB x 0.2

With Resume being calculated as follows:

Top 25 wins: 16 points
other top 50 wins: 13 points
other top 100 wins: 5 points
sub 100 losses: -1 point
sub 200 losses: -5 points

The original T-Ranketology got a score of 312 on bracketmatrix, by getting 65 teams right, nailing the seed on 31 and within 1 seed on 24 others.

This version of T-Ranketology gets a score of 351 (which would tie for first place this year), by getting 66 teams right, nailing the seed on 44 teams, and within 1 seed on 21 others. This algorithm gets Tulsa and Vanderbilt into the tournament, but leaves Providence and Wichita State as 2nd and 3d teams out, respectively. (St. Bonaventure is the first team out.) Saint Mary's remains in the field, though as a 10-seed instead of an 8-seed. Florida also sneaks into the tournament.

If you want to see the bracket produced by this algorithm, it's here.

Clearly, this version of the algorithm is "over-fit" to this year's results. But I think this exercise does provide some insights. Most obviously, the "resume" rank is extremely important. This is how you get Tulsa into the tournament. You have to really value those wins against "top 25" and "top 50" teams. Bad losses don't matter too much, at least not much more than they already matter for the other ranks. The elo and WAB ratings add a little, but not much, to the analysis.

So, this is the algorithm I'll go with next year. The committee will probably do something entirely different and prove, once again, that you can't model Madness.

Tuesday, March 15, 2016

Thoughts on the bracket, part 2: The Snubs

The T-Ranketology algorithm had three teams in the tournament that the selection committee found underserving: St. Bonaventure, Saint Mary's, and South Carolina.

It's pretty obvious what they have in common: they all start with the letter "S". Frankly, anti-S discrimination is as good a theory of why the committee does what it does any other. But let's dig a little deeper into their resumes, and that of the other cause célèbre, Monmouth.

Monmouth

T-Ranketology was not surprised by Monmouth's exclusion, as they were the 12th team out according to the algorithm. This is because they got killed in the "resume" column because of their three bad losses to sub-200 teams.

Monmouth is a tough case because they had the four great wins in the non-conference, against UCLA, USC, Notre Dame, and Georgetown -- all away from home. Since high-major teams have no incentive to play true mids or low-majors on the road, the only path for a team like Monmouth to an at-large bid is go giant slaying on the road, and that's exactly what they did.

Unfortunately, the UCLA and Georgetown wins ended up not looking so great in the committee's eyes, because Georgetown was a sub-100 RPI team and UCLA was 99th. (Indeed, if they'd lost to Georgetown that would counted as a "bad loss"!) This is true even though those were both true road games, which makes them impressive wins by any measure, except any measure the committee pays attention to.

In the end, it was the three sub-200 losses that killed them. I'm pretty sure no team has ever gotten an at-large bid with three losses in that category. This is somewhat unfair to Monmouth, of course, because most at-large contenders do not play very many sub-200 teams on the road. As a result, we don't have an intuitive feel for how often at-large contenders should really lose these games. Monmouth played 11, and they went 8-3. That's not good, but how bad is it?

Easy answer: too bad. I think Monmouth is in the tournament if they go 9-2 in those games. But it was three strikes you're out.

South Carolina

South Carolina was the last team in according to T-Ranketology, and no one is weeping over their omission from the field. They played a crappy schedule, which got them out to a 14-0 start. They were even 20-3, but went just 3-6 in their last nine games. Although their schedule was weak, they actually performed admirably against it, compiling +2.0 WAB, which means they won two more games than you'd expect an average bubble team to win. But (as we'll see) the committee comes down hard on teams that didn't "challenge" themselves during the non-conference, and South Carolina's OOC SOS was 271st according to the RPI.

Since South Carolina is a major conference team that has no excuse for playing such a weak schedule, and they were right on the bubble by all metrics, no one cares about saying sayonara to South Carolina. Hmm, maybe I should write a song called "Sayonara, South Carolina."

Saint Mary's

In my opinion, Saint Mary's was the real snub this year. T-Ranketology had them into the field easily as an 8-seed, and they were in a majority of final brackets at bracketmatrix.com. They had a weak nonconference schedule, but Seth Burn has already detailed how well they performed against the schedule they played, in terms of Wins Against Bubble. They also did well in other metrics traditionally associated with good tourney resumes, such as elo.

Ultimately, they were done in by their lack of "good wins." Their record against the top 100 was great -- 6-3, but the committee cares most about number of wins, not winning percentage. And in the all important "wins against top 50" they had just two. And both those were against Gonzaga, which only squeaked into the top 50 after they beat Saint Mary's in the WCC championship game.

This is a case, I think, where the committee was forced to come face-to-face with the absurdity of its own metrics. Heading into the game against Gonzaga, Saint Mary's had a blank resume, highlighted by zero top-50 wins. They lost that game, convincingly. But now, because of that loss, they had an infinitely better resume, with two top-50 wins. How could a loss possibly improve their resume so much?!

That's a little thing called cognitive dissonance, Jack.

The committee did what anyone does when experiencing cognitive dissonance: it moved on as quickly as possible. Buh-bye, Saint Mary's.

St. Bonaventure

The Bonnies were the last of the S-nubs, and the most surprising, as they were in most every final bracket. But upon examination, their big calling card was a high RPI. The rest of their resume metrics were bubblicious, or worse: 49th in elo, 50th in "resume" (good wins minus bad losses), and 69th in WAB. I shed no tears for St. Bonaventure. Indeed, their exclusion is another sign that raw RPI is (appropriately) not much of an independent factor in the deliberations.




Monday, March 14, 2016

Thoughts on the bracket selection, part 1

I followed the "bracketology" debate more closely than usual this year, for two reasons: (1) until recently, the Badgers were a bubble team (at best), and (2) I used the power of T-Rank to produce an objective bracket prediction (T-Ranketology).

I was satisfied with T-Ranketology's results. It missed three at-larges: Syracuse, Vanderbilt, and Tulsa (instead of Saint Mary's, St. Bonventure, and South Carolina). This was about average for the brackets over at bracketmatrix.com, and if you look at the final "consensus" bracket, T-Ranketology would have had 67 of 68 (with only South Carolina not being a consensus tourney team). Seeding predictions were pretty good, too. All in all, not bad for an algorithm.

There's a lot of hootin' and hollerin' about the committee's selections, and I agree with much of it. But I wanted to take a look at the six teams T-Ranketology was wrong on, plus one other, to see maybe what the committee was thinking and if the algorithm could be improved to reflect that thinking or lack thereof.

Syracuse

Syracuse's overall resume was not great, but the one big trump card they had was a road win at Duke, which is worth like a million resume points. Add a couple other decent (early) wins, and they're left ranked No. 32 in T-Ranketology's "resume" ranking, which is an attempt to simulate the committee's fixation with "top 50" wins, etc. They were also 50th in T-Rank's "wins against bubble" rating, so right on the bubble there. But they were very poor in basic RPI, and had tanked coming down the stretch, leading to a bad elo ranking (which usually pretty well approximates current "sentiments" about how good a team is). All in all, this left Syracuse the sixth team out in T-Ranketology.

But I'm not going to worry about fiddling with T-Ranketology to get Syracuse into the field, because they were clearly a special case.
The committee chairman had been saying for weeks that Syracuse was going to get special consideration because Syracuse played like shit when Jim Boeheim was serving his 9-game suspension in the middle of the season. Basketball people can look at that argument and see the absurdity on its face -- Syracuse lost five games without Boeheim, and even if you think he is super-god-coach, they almost certainly lose 4 if not all 5 of those games with him. But the committee is not composed of basketball people, so alas. In any event, the handwriting had been on the wall: barring a monumental collapse (which almost happened) Syracuse was going to be in the tournament. If I had been manually fiddling with the T-Ranketolgy bracket, I'd have put Syracuse in for sure.

Vanderbilt

Vanderbilt was the eighth team out in the final T-Ranketology, mainly because it was not in the top 50 of any of the four resume-based metrics the algorithm considered. On the other hand, unlike Syracuse, it was not terrible in any of those metrics, finishing top-70 in all of them.

Clearly Vanderbilt got in because of its efficiency rating. Vandy finished the season ranked 27th in the Kenpom ratings, and top-30 teams always get in nowadays. I think this is why they're slotted in the play-in round against another Kenpom darling, Wichita St.: the committee put two good teams with bad resumes into the ring against each other and said, "If you're so good, prove it."

When I first started T-Ranketology it was loosely modeled after the Easy Bubble Solver, which just averages RPI and Kenpom ranking. Of course, I used T-Rank instead of Kenpom rating, but it was pretty much the same idea. Eventually I took out the efficiency rating component because for good or bad I think it just doesn't play a very big role in the selection or even seeding process, at least not systematically. But I think there's good evidence that it comes into play in edge cases, and I think it's pretty clear that's what got Vandy in. I may have to work in some kind of "top 30" efficiency rating bonus to account for this.

(By the way, I think Syracuse was also helped by a decent Kenpom rating, though I don't think it wouldn't have been enough without The Boeheim Excuse.)

Tulsa

Tulsa was the real shocker, what John Gasaway calls the committee's annual "grenade." But it wasn't a huge shocker to T-Ranketology, which had Tulsa just the fifth team out -- ahead of both Syracuse and Vanderbilt! Indeed, before its loss to Memphis in the AAC tourney on Friday, T-Ranketology had Tulsa the last team in the field.

Why? Like Syracuse, Tulsa scored well in the "resume" score that approximates top 50 wins, etc. This is by far the stupidest possible measure you could come up with, but I'm pleased to say that I think I've done a pretty good job of modeling this particular madness. Tulsa's inclusion in the field of 68 shows that I probably need to weigh it even a little more.

But I think there's another lesson for Tulsa's selection. The committee first gets together early in championship week to get ahead start on the process. As a result, by Friday (when Tulsa got stomped by lowly Memphis) the committee has already made some provisional decisions about which teams it thinks are good enough. I'm pretty sure that Tulsa's resume had already been found deserving by Friday. Here's the way the human mind works: once it decides something, that decision sets. It's like a boulder in a divot; you need a big shove to get it moving again. When new information comes in, we don't start at square one and reevaluate the decision with a blank slate. We say: is this new information enough of a big deal to make me go through this whole process of deciding again? Unsurprisingly, the answer is usually no. This fundamental quality of human psychology (call it laziness if you wish) got Tulsa into the tourney.

(This, by the way, is why college-football-playoff-style in-season tourney rankings are a terrible, terrible idea.)

Well, this has gotten tl;dr so I'm going to stop now. I'll try to post later today with my profound insights into the "snubs."