Showing posts with label WAR. Show all posts
Showing posts with label WAR. Show all posts

23 May 2014

A WAR for Triple Crowns

The Triple Crown is rarely bestowed upon any hitter in Major League Baseball. It is also a title that generates a great deal of discussion about its value.  Traditionalists, and perhaps mainstream fans, see the achievement as reaching a pinnacle feat of greatness.  You will at times hear about Frank Robinson's Triple Crown if you stick around at Camden Yards during a rain delay.  The beauty of the recognition of this feat is largely a product of cricket. 

Henry Chadwick, who was quite familiar and fond of cricket, was looking for a way back in the 1800s to easily communicate the events that transpired during a game.  He leaned on the simple metrics used in cricket to credit the players.  Simply put, the forefather of traditional statistics is cricket.  In that way, you can imagine that perhaps these metrics are better at describing that sport as opposed to baseball. 

Regardless, one would imagine that over the following 150 years that better metrics would be created.  Many would argue there are better metrics.  However, the simplicity in communicating the game that Chadwick devised has stuck around and has largely impacted our perception of this game.  This traditionalism vs. modernism fight has been largely on display on issues surrounding these metrics and the Triple Crown is not foreign to those discussions.

For example, this past winter I encouraged an aspiring writer who sent me a column on Triple Crowns to seek out writing opportunities at other sites.  I was simply not interested in how he was trying to place different types of importance on those feats.  The problem I really had with the article was that it was not really trying to bridge a gap between the metrics struggle.  For instance, why would home runs and runs batted in be equivalent in value?  That would shift importance towards that component and suggest that the three variables are not equal.  I think that runs opposed to the core belief in the award as well as abuses the concept of combining data to say more about a set of data.

That column has stuck in my mind though.  Although we tend to write about talent and future value, there is a place for describing what really happened.  Life is not always what will happen or how past events inform us about what will happen.  Sometimes, life is simply about what occurred and sometimes what occurred in a subset of data.  With that in mind, I can appreciate the Triple Crown.  No, it does not say this individual or that individual is the best player in history, but it is a rough award that notes players who were exceptional.  Of the 14 players who won the Triple Crown, maybe Tip O'Neill was the worst and he played in those weird days of baseball in the 1880s.

Anyway, I decided to tinker around and put the Triple Crown on the same scale as fWAR.  Based on that distribution, I devised formulas that would convert batting average, home runs, and RBIs on a similar scale.  One aspect we can hem and haw on, I took qualified players and then treated home runs and RBIs as a function of play appearances.  That felt right to me, but it certainly is another level of data manipulation for purists of cricket-derived statistics to rail against. 
avgWAR = 61.021 * AVG - 14.015
hr%WAR = -1.778 * LN(PA/HR) + 8.7967
rbi%WAR = -4.905 * LN(PA/RBI) + 12.965

Anyway, I then averaged the three WAR scores to devise something I called tcWAR.

Below are the top ten players from 2010 through 2013, career numbers:




Name Team tcWAR PA HR RBI AVG
Miguel Cabrera Tigers 5.0 2685 156 507 .337
Ryan Braun Brewers 4.2 2244 108 364 .316
Adrian Beltre - - - 4.2 2510 126 401 .314
Carlos Gonzalez Rockies 4.2 2193 108 364 .311
Troy Tulowitzki Rockies 4.1 1850 90 309 .307
Robinson Cano Yankees 4.0 2755 117 428 .312
David Ortiz Red Sox 4.0 2194 114 361 .300
Allen Craig Cardinals 4.0 1420 50 247 .306
Josh Hamilton - - - 4.0 2381 121 401 .296
Joey Votto Reds 3.8 2568 104 345 .317




How is this year's group of batters projected to do (statistics through May 21st)?



Name Team tcWAR PA HR RBI AVG
Troy Tulowitzki Rockies 6.0 183 13 35 .378
Yasiel Puig Dodgers 5.0 187 10 37 .333
Charlie Blackmon Rockies 4.7 185 9 32 .335
Miguel Cabrera Tigers 4.7 183 7 40 .321
Victor Martinez Tigers 4.6 180 12 28 .329
Giancarlo Stanton Marlins 4.6 205 12 44 .305
Justin Morneau Rockies 4.6 174 9 32 .321
Brandon Moss Athletics 4.6 178 10 40 .301
Nelson Cruz Orioles 4.3 187 14 41 .282
Michael Brantley Indians 4.2 192 9 36 .302



What should we take from this?  Well, very good players tend to do very well when it comes to these categories.  The absence of Jose Abreu also shows the uniqueness of MLB's current home run leader.  Also, Nelson Cruz is proving to be one of the most meaningful free agents of the past year, showing that having a haphazard recruitment process sometimes actually can work though it probably does not speak highly for the process itself.

I think that might confuse some into thinking that doing well in these categories means that these metrics are incredibly useful in determining which will do well as opposed to those who performed in the opportunities that were given to them.  For instance, you may be able to perform insanely well at a Wall Street Firm, but your online law degree probably will not provide you much of an opportunity to test yourself in that venue.  In other words, equally able players may not be given the same chances to perform.  Second, these metrics do not truly define baseball.  It is a much more varied game that cannot be reduced to only three considered counting statistics.  In other words, you may be great at sprinting while being pretty awful at hurdles.  Both have value, but if you only look at sprinters then you are missing a lot of what it means to be good at track events.

01 March 2012

Comparing fWAR with rWAR

Here is just a short post today.  People often think of rWAR and fWAR as being equal to each other because they are both trying to determine the overall value of a player's performance.  They use different means, but try to get to the same place.  However, this does not entirely make sense to me because the assumption is that the two metric would have the same statistical spread.  I decided to show them simply side by side below using statistics from 2011.

Pitching

Below are the WARs for pitchers who qualified for the ERA title.


These two for the most part match up well.  rWAR is a bit more extreme on the ends with fWAR showing a slight bump throughout the middle, particularly with players below a WAR of 2.  fWAR has a mean of 3.2 while rWAR is 2.9, which amounts to a total difference of about 20 WAR between the two statistics.  As populations, they do not appear to be significantly different (p=0.40).

Position Players

Below are the WARs for position players who qualified for the batting title.


Again, we see something similar to the WAR for pitchers where at the extreme ends, rWAR gives greater positive values and lesser negative values while fWAR gives higher values throughout the middle.  The differences between these two populations approach significance (p=0.10).  The average WAR was 3.2 for fWAR and 2.8 for rWAR.  The difference in total WAR was 60 WAR advantage to fWAR.

Conclusion
I think the take home message here is that while you may not be making a grand mistake by doing something like adding them together and dividing in half, you certainly should not think the two statistics are equivalent in magnitude.

11 May 2011

Are All Divisions Created Equal? Team WAR (2002-2010)

Many a fan, particularly in Baltimore, has uttered the words: if only we played in another division.  Without a doubt, the AL East is a difficult division to play in.  The Yankees and Red Sox have considerable resources that help them sign top talent in the offseason and gives them enough of a margin of error to absorb bad contracts.  The Rays are saddled with severe cash restrictions, but their front office finds remarkable ways to remain competitive.  Finally, the Blue Jays are a club that always appears to be underrated.  It is tough to be an Oriole fan at times with all of these successful teams to play against with an unbalanced schedule.  This leaves many wondering what would it be like if the Orioles were to play in a different division and makes me want to try to quantify it.

This is not exactly an original idea.  It has been addressed before by Stacey Long over at Camden Chat.  They decided to use the Orioles' winning percentage in each division to determine how the team would perform with various unbalanced schedules.  In that article, Long found that the Orioles' record in 2009 would have improved by four games had they played in an easier division.  I think that is solid work, but I wonder if the number of games the Orioles play against other teams in the other divisions is robust enough.  Daniel Moroz of Camden Crazies wrote over at Beyond the Box Score a short piece trying to determine what the difference playing between divisions would be.  Moroz used an elegantly simple set of assumptions (e.g., typical wins by AL East teams, an 81 win average for AL teams, interleague play winning percentages) along with Bill James' log5 calculations.  He found that unbalanced schedules could cause changes of a win.  That is not much to be concerned about.  So, two solid articles with two different conclusions.

This well tread idea needs, perhaps, another way to address it.  I propose that WAR (Wins Above Replacement) could be used to determine how many games a team could win by shifting divisions.  You may have noted in the past that if you sum up all of the WAR of individuals on a single team, you wind up with a number that is anywhere from 35-55 wins less than the actual total wins that the ball club earned.  This makes sense as a replacement level team (roughly defined as a team composed of AAA players) would be able to win a game here or there.  WAR actually relates rather well to actual wins.  The graph below illustrates how WAR relates to actual wins for AL teams from 2002 to 2010.

click on image to make larger


The trend line shows a good fit between wins and WAR with an R-squared of 0.81.  The y-intercept (where the line crosses the y axis (the vertical one)) is the number of wins that would be expected to be won by a replacement level team.  As this can be done with all of these data points from 2002 to 2010, we can also do something similar for teams in each division for each year.  With fewer data points, the correlation will not be as strong, but using replacement level performance as a base line for each division would provide a different way to measure how a team would perform in a different division. The following graph illustrates Replacement Level Wins (essentially Average Wins minus Average WAR) for each division.

click on image to make larger

The graph above passes a general smell test for me.  My perception has been that the AL East is the toughest division to play in because there is a concentration of talent of both young players (Tampa Bay Rays) and free agents (New York Yankees).  It has also been my perception that this concentration of talent has been challenged at times, which is also indicated in the graph above.  What I find interesting is that over this stretch of time, a replacement level team shifting from the AL East to the AL Central or AL West would improve by one and three wins, respectively.  If you translate that into free agent money (4.5 MM for every win), you could say that it costs 4.5 MM less to compete in the AL Central or 13.5 MM less to compete in the AL West.

Again assuming that the replacement level is additive, we can determine how well the Orioles would perform in "weakest" AL Division each year over the past nine years.


Comparing our results to what Long and Moroz found, we see in general that our numbers fit in more with Long's in the idea that shifting divisions would cause a decent sized shift in the number of wins.  When we specifically unravel Long's predictions for the Orioles in the AL Central from 2002 to 2009, we do not see a strong match between our efforts.  Years in which I show little difference between the AL East and AL Central (e.g. 2005), she shows a difference.  In my opinion, I think this is largely a product of there may not being enough games played when using one team's winning percentages to determine division strength.  What is certainly known by all of these efforts is that simply moving the Orioles to another AL division is not a cure for what ails them.

In the near future, I will be using this approach to address the addition of a fifth team in each league's playoffs and will try to determine how good that team would be in comparison to the current field.