Showing posts with label Fangraphs. Show all posts
Showing posts with label Fangraphs. Show all posts

23 August 2018

A Look At The Distribution Of Prospect Values

Over the past few years, there’s been a movement to grade prospects by not only overall rank and organization rank, but also by future value. Future Value is not very well defined, but per MLB is a 20-80 scale in which a ranking of 20-30 is well below average, 40 is below average, 50 is average, 60 is above average and 70-80 is well above average. A player with a 65 FV is someone who could develop into a future impact Major Leaguer, perhaps an All-Star-caliber standout. Fangraphs defines future value, using a similar grading model and provides a chart showing how future value can be defined using WAR.



In general, I’m not a fan of using such definitions as Top #2-3 player or #2/#3 starter because it’s subjective and can’t be easily determined and prefer using WAR values because that at least is easily definable and somewhat objective. WAR may undervalue a strong player that is hurt for a few years, but any clear definition will have its limitations.

In order to learn about future value, I decided to build a metric called actual value. Future value attempts to predict how much production a prospect will produce using WAR, and actual value attempts to measure how much production a prospect actually did produce. In essence, I defined actual value in such a way that future value attempts to predict actual value. 

To do this, I looked at players from 1990-2016 and measured each prospects’ production while they were under team control or until they had six or more years of service time. Once players have six or more years of service time, they are eligible for free agency. Then, I determined their rookie year and measured their average value per year or actual value using the following scale.



With this distribution of actual value illustrating how prospects performed, we can understand what a reasonable distribution for future value might look like for the future. I looked at the performance of prospects over two time periods, 1990-2016 and 2004-2013 defined by a prospects rookie year. Here were the valuations this process returned from 1990-2016.



Each year from 1990-2016, there were roughly 18.5 prospects that had an actual value per year of 2 WAR or more. This may seem surprising at first glance, but one should consider that from 2000-2017, there were roughly 130 offensive players and 90 pitchers that were worth more than 2 wins or above average. This includes players that were one year wonders like 2009 Jason Kubel (2.5 fWAR) or 2016 Michael A Taylor (3.2 fWAR). They may have been worth 2 fWAR in a single season, but weren’t above average players over an extended period of time. When looking at players that have staying power, there are significantly fewer than 220 at any given time – maybe around 150 total. If there are few average and above average players in the league, then there can be few prospects each year to replace them. If 20 prospects each year are above average, and there are 150 players total above average, then the replacement period would be about 7.5 years. This seems to suggest that 20 average prospects per year is within reason.

There were significantly more prospects each year that ended up being below average and the vast majority were either 25s or 30s. This makes a lot of sense because teams have injuries and need to call up their prospects to make it through the season. In addition, teams would prefer to give their prospects chances to be successful because they’re cheaper than free agents. However, the vast majority of prospects either fail or struggle to be better than replacement level.

The story is reasonably similar when looking at prospects that became rookies from 2004-2013. There were about 20 rookies per year that because average or better players in the majors. Another 22.5 per year were below average, but still worth over 1 WAR per season. The vast majority of rookies, or roughly 184 per year, ended up being worth less than 1 WAR per season or clearly below average. Both sets of data tell the same story: being successful in major league baseball is difficult.



If only twenty players per year have above average actual value, than prospects ranked outside of the top 100 should not be expected to have a high future value. After all, future value is trying to predict and should have similar behavior to actual value. Prospect lists are tricky because they grade all prospects and not just those that should be expected to enter the league each year. So, I would argue that it’s reasonable to presume that all prospects considered to be at least average should be in the top 100, and likely all prospects with a future value of 45. I would say that all prospects with a future value of 40 should be at least in the top 125 prospects and all prospects with a future value of 35 should be in the top 200 prospects. Players ranked worse than 200 should be seen as slightly better than replacement players at best and ideally will be either bench players, middle relief or minor league depth pieces.

Fangraphs provides a board with its top 850 prospects as ranked at the beginning of the season with both future value grades and estimated time of arrival. We can use this board to determine the time period that their prospect list measures as well as see whether their distribution seems reasonable.



If they still use the future value definitions defined by McDaniel, than their distribution is difficult to comprehend. In their defense, it appears that they consider their FV to be their peak level rather than their average level. In addition, they deserve some leeway for injuries. Still, it is highly unlikely that there were will be 54 prospects in the 2018 rookie class that end up being average or better or 110 prospects that are worth 1 fWAR or more. It seems that a 55 in the Fangraphs distribution is the same as a 50-55 in the historical distribution, a 50 in Fangraphs is the same as a 35-45 in the historical distribution, a 45 in Fangraphs is the same as a 30 in the historical distribution and a 40 in Fangraphs matching a 25 in the historical dataset. Perhaps this is why Fangraphs only gives valuations for players that are ranked up to 45?

Even prospects graded as a 30 have a surprisingly high amount of value using a WAR linear model. These players typically produce about .2 fWAR per year and last for about 4.5 years. Using our predicted arbitration rates would suggest that these players produce $9M in value (at $10M per win) and cost $2.5M for a total surplus value of $6.5M. This seems surprising at first glance, but consider that players like Caleb Joseph and Ryan Flaherty rate as having an actual value of 30. These players aren’t great, but Joseph is certainly a decent backup catcher option while Flaherty was a decent utility player. Despite the fact that these players are valuable, it is impossible to build an average team with players just like those two, explaining why teams are willing to trade them away for marginal improvements. Treating WAR as a linear model does have its limitations.

Players ranking 35, 40 and 45 in actual value have a preposterous amount of value using this metric. 35 actual value level players produce roughly 4.2 WAR of production over their six years of team control or produce $42M in value (at $10M per WAR) but only cost about $10.6M putting their value at $30M. 40 level players produce roughly 7.2 WAR of production over their six years of team control or $72M in value but only cost $15.6M putting their value at $56M. 45 level players produce roughly 10 WAR of production over their six years of team control, but cost roughly $24M putting their value at $76M. Presumably, Orioles fans would not have been happy if the Orioles received just one 45 level player (top 60-90 prospect) in return for Machado.

Part of the reason why these players have such high values is that non-arbitration player salaries are unfairly low. If a win is valued at $10M, and the salary for a non-arb player is $550k, then it is relatively easy for a non-arb player to be underpaid. A reasonable minimum salary of $1.5M would rectify the situation where near replacement players have significant value. Part of the reason is that having a slightly above replacement player play instead of a below replacement player can have significant value. The fact is that teams are willing to pay a premium to avoid having to use a player like David Hess for 65 innings.

Future value may be the attempt to determine actual value, but that doesn’t mean they are the same. There’s a significant amount of uncertainty when predicting future value, and the distribution of prospect performances suggest that it is far easier to overestimate a prospect than underestimate a prospect. Most prospects that make it to the majors aren’t successful, and all prospects that don’t make it to the majors aren’t successful by definition. This uncertainty will mean that a player with an actual value of 40 (as determined in hindsight) is far more valuable than a prospect with a future value of 40. In a player with an actual value of 40 is worth $56M, than a prospect with a future value of 40 is worth significantly less. This is one reason why prospects aren’t valued as actual value suggests. In contrast, measuring the value of top 100 prospects take failed prospects into account and therefore provide a reasonable baseline for value.

In addition, a team filled with 35 AV players that are under team control, may be fiscally successful, but will also be terrible. In some roles, a 30 or 35 AV player can be acceptable, but in other roles (starters) they’re not particularly ideal. Obviously, it’s important to be fiscally responsible, but teams are in the business of maximizing their win total and not their excess value total. Others believe that teams would rather win now rather than later, and therefore future wins should be discounted.

Going forward, looking at how many prospects historically fit into each category of actual value should help analysts build future lists using future value. It doesn’t make sense to predict that hundreds of prospects will have a future value of 45 if few players ultimately have an actual value of 45. Using historical data to accurately assess the value of prospects will lower these rankings and make it more difficult to excite casual fans about their teams’ future. Still, even if it’s bad press, it makes sense to admit the reality of prospect value. Having a better idea how prospects historically perform can help teams plan for their future, and ultimately fans don’t really want to see how their prospects develop but rather watch their team be successful at the major league level.

09 February 2017

A Look At Fangraphs Projections

Certain things always happen during the offseason.  Free agents are signed, players are traded for, tons of digital ink is spilled discussing all the possible rumors and Fangraphs predicts that the Orioles will be the worst team in the AL East. At this point, it probably isn’t very surprising to learn that Fangraphs expects the Orioles to win ten fewer games in 2017 than they did in 2016. It’s probably more surprising to learn that Fangraphs does project the Rangers to win twelve fewer games in 2017 than they did in 2016. But how accurate is Fangraphs exactly? They certainly missed on the Orioles last year, but did they do better predicting other teams’ results?

Fangraphs had a surprisingly good year from a wins perspective in 2016. They were off by roughly 5.6 wins per team in 2016. In contrast, presuming that teams that every team would win 81 games would have been off by slightly more than 9 wins per team making it roughly a 40% improvement than just picking each team to win the same amount of games.

While they successfully predicted the record of 11 teams within 3 wins of their 2016 results, they also were off by 9 or more wins for 6 teams. Interestingly, Fangraphs did a worse job projecting five other teams than the Orioles in 2016. Presuming that each team would go .500 would have resulted in being off by 9 or more wins for 14 teams, while being within 3 wins for only 6 teams. This indicates that while Fangraphs had poor predictions for a number of MLB teams, using these predictions is better than using nothing at all. It also shows that Fangraphs should be expected to have poor predictions for roughly 20% of MLB teams.

Fangraphs did a decent job predicting runs scored for each team as they were off on average by .266 runs per game. If someone knew that teams would average roughly 4.48 runs per game, and predicted that each team would score that amount, then they would have been off by roughly .289 runs per game. Given that it’s hard to know before the season how many runs will be scored, this shows that Fangraphs had some success at predicting runs scored. Fangraphs did better when it came to runs allowed. On average, Fangraphs was off by roughly .3 runs per game. If one presumed that each team would allow 4.48 runs per game, then presuming each team would allow the same amount of runs means that the average prediction would have been off by .36 runs per game. Clearly, using Fangraphs projections have some value.

Fangraphs does worse when trying to project earned and unearned runs. Fangraphs was off by an average of 59 earned runs per team. If one used the actual number of runs scored, than presuming that each team allowed the same amount of earned runs would have resulted in the projection being off by 51.4 runs. If one used the number of earned runs projected to be scored by Fangraphs, then the average team would have been off by 69.6 earned runs. This isn’t an impressive result.

Likewise, the Fangraphs projections predicted that each team would allow 75 unearned runs. In reality, teams only allowed 52.76 runs. This is a pretty big difference. On average, Fangraphs was off by 23.7 unearned runs per team. If one presumed that each team gave up an equal amount of the actual unearned runs allowed, then each team would have been off by 10.4 unearned runs. If one used the number Fangraphs projected, then each team would have been off by 22.9 unearned runs. This indicates that Fangraphs is unable to project unearned runs. It also suggests that Fangraphs is unable to successfully predict the number of runs allowed in a season, but has the ability to guess which teams will allow the most or least runs only in context. This isn’t as useful as how many runs a team will allow.

The Fangraphs projections undoubtedly have some predictive ability and are better than using nothing. Sometimes, that can be very valuable. For example, suppose someone builds a model that can successfully predict whether a stock on Wall Street will go up in value that works 5% more accurately than the current models. This model may only have minimal predictive ability but is still good enough to be worth billions of dollars. A small increase in predictive ability can be highly valuable.

However, it doesn’t change the fact that a small increase in predictive ability is still only a small increase. Sure, it tells us more than we knew before, but that’s not the same as calling it gospel. The model in the paragraph above will fail a large percentage of the time. Such a model can both have extreme value, but still be wrong a significant amount of the time.

What does this mean for Orioles’ fans? The Fangraphs projections do have some validity to them, but they’ve historically missed on a large percentage of teams. Their results should be taken seriously, but with the understanding that they aren't holy writ. They’re neither perfect nor worthless. The same is probably true for PECOTA. They’re better than nothing, but that doesn't make them great.

24 May 2016

Looking Back At Fangraphs' Projections

After a fourth of the season is in the books, the Orioles are 26-16 and have the second highest winning percentage in the majors. Naturally, this vindicates Fangraphs which projected that the Orioles would win the AL East this year. Wait, what’s that you say? Oh yeah, they projected the Orioles to be in last place and not first. On the bright side, at least they didn’t project the Orioles to win 72 or 73 games. I guess the obvious question to ask is what happened.

For starters, it’s worth noting that the Orioles aren’t their largest miss. When comparing their actual winning percentage to their projected winning percentage, Fangraphs has been worse when projecting the Twins (off by 21.9%), Phillies (off by 17.3%), Braves (off by 14.1%) and Astros (off by 15.9%). Meanwhile, Fangraphs projected the Orioles to have a 49.4% winning percentage and the Orioles actually have a 61.9% winning percentage. On average, Fangraphs is off by 6.77% or roughly 11 wins per team over a full season. If someone had projected that each team would go .500, then they’d be off by 7.73% or 12.5 wins per team over a full season. So far, Fangraphs has been more accurate than just presuming that each team would go .500, but not by very much. It is still early in the season though, but at the current moment, well.



As for the Orioles specifically, Fangraphs projected that the Orioles would score 4.64 runs per game and the Orioles have actually scored 4.55. The Orioles are on pace to score 15 runs fewer than Fangraphs projected. This is reasonably close and suggests that Fangraphs overestimated the Orioles offense. These results illustrate that the problem with their projection wasn’t runs scored. Rather, their problem was with runs allowed. Fangraphs projected that the Orioles would allow 4.69 runs per game. The Orioles have actually allowed 3.98 runs per game. The Orioles are on pace to allow 116 runs less than their March projections and 83 runs less than their current projections. If the Orioles can keep that up, then they’ll prove the computers wrong.

The reason for the discrepancy obviously comes on the pitching/defense side of things. But is it all of the pitching? Fangraphs originally projected that the Orioles starting pitching would pitch 935 innings, with a 4.38 ERA and thus give up 455 earned runs. So far, they’re actually doing pretty well with this prediction. The Orioles starting pitching is on pace to throw 914 innings with a 4.44 ERA and to give up 451 earned runs. The Orioles starting rotation would be on pace to allow 461 earned runs if it threw 935 innings. It’s pretty clear that Fangraphs has nailed the Orioles’ starting pitching so far.

The problem is that Fangraphs also projected that the Orioles bullpen would pitch 523 innings, with a 3.76 ERA and give up 218 runs. So far, the bullpen is on pace to pitch 521 innings, but with a 2.67 ERA and thus give up 154 runs. In addition, Fangraphs projected the Orioles to allow 86 unearned runs. So far, the Orioles have allowed 10 and are on pace to allow just 39. Back in March, I wrote an article discussing how Fangraphs projected standings is likely flawed due to how it accounts for unearned runs. It seems like that flaw has come back to bite them.

This flaw has a surprisingly large impact. Using a t-test, Fangraphs' projected runs allowed on the team level is statistically different then the actual results so far (t>.9781). But Fangraphs' projected runs allowed and actual runs allowed by the starting rotation isn't statistically different (t>.1410). This is also the case for the bullpen (t>.2019) suggesting that they're doing a reasonably good job projecting starter and reliever performance. But a t-test comparing the amount of total unearned runs allowed by team to the projected number of unearned runs allowed by team is statistically different at the <.0001 level. Their inability to predict unearned runs significantly weakens the value of their runs allowed projections. Fangraphs' results in this regard are so poor that the Orioles aren't the most egregious case. The Royals, Indians and Rays are each on pace to outperform their projection by over 50 runs. They were projected on average to give up 77 unearned runs and are on pace to allow 21 each. That can't be good.

Using the Wayback Machine, since Fangraphs doesn’t save its original depth charts, it’s possible to review Fangraphs’ projections on March 11th, 2016. So far, a number of crucial pitchers in the bullpen have outperformed Fangraphs’ assumptions. For example, Fangraphs thought that Zach Britton would be good, and projected him to allow 19 earned runs while throwing 65 innings. Britton is on pace to allow 13 runs while throwing 76.66 innings. But Brach is the real lynchpin. Fangraphs projected Brach to give up 21.75 runs over 55 innings. Brach is on pace to give up 12.7 runs over 98.5 innings. That’s a 26 run swing right there, presuming he throws 98 innings in relief. Givens is on pace to give up 13 fewer runs than projected. McFarland, Bundy and Matusz are the only underperforming relievers and they’ve thrown 37 of the bullpens’ 135 innings. After those three, the reliever with the worst results is Darren O’Day with his 2.76 ERA.

Furthermore, the Orioles four elite relievers are throwing 60% of all the innings thrown by the bullpen despite being projected to throw 44% of the innings. Bullpens have better ERAs when their best pitchers throw more innings. That, combined with unexpectedly strong performances from Worley (0 ER in 14 innings as a reliever) and Wilson (1 ER in 8 innings as a reliever) has meant that the Orioles’ bullpen has vastly over performed.

On the defensive side, the Orioles may rank poorly in Fangraphs’ Def stat (20th in MLB), but they only allowed 18 errors in their first 42 games. The Orioles allow .55 unearned runs per error which is worse than league average. But they have the fifth lowest error per game ratio in MLB and is why the Orioles have given up so few unearned runs. On a simplistic level, if the Orioles continue to give up few errors, they’ll allow few unearned runs. On a more complex level, does this mean that the Orioles defense is better than it appears? The Orioles fielding weakness is corner outfield defense, which UZR may not be able to measure properly. Their BABIP is league average, but it would take an analysis to determine the type of batted-ball contact that their pitching has allowed. Further, Inside Edge noted in a blog post that the Orioles rank tenth in defensive giveaways.

The Orioles have outperformed Fangraphs expectations so far due to their excellent bullpen and their ability to avoid unearned runs. They have been slightly lucky, but it makes sense that a team with a strong bullpen will get lucky. Going forward, perhaps this suggests that the Orioles shouldn’t be focusing on adding starting pitching, but perhaps adding a good reliever to ensure the bullpen can continue performing at its current level. It’s pretty clear that the Orioles’ plan is to hope that their offense will maintain its current standard and that their bullpen can continue dominating.

17 March 2016

Two Problems With Fangraphs' Projected Standings

A few months ago, Fangraphs published its projected win totals for 2016. They weren’t very complimentary of the Orioles as they projected them to win just 80 games and be the worst team in the AL East. The Royals, the 2015 World Series Champions, are projected to win only 77 games this season despite minimal changes. I thought it would be worth looking into why Fangraphs thought that the Orioles might struggle.

Fangraphs projects the Orioles to score 4.64 runs per game (752 total) while allowing 4.69 (760 total). If the Orioles did score 752 runs, then that would be the most they’ve scored since 2008. It's clear that the projections predict that the Orioles’ pitching is their weak point.

Fangraphs WAR depth charts tell a similar story.  They project the Orioles’ to have the eighth highest offensive WAR totals, but tied for the sixth lowest pitching WAR totals. However, what’s interesting about the Fangraphs WAR depth charts is that they project a total of 1085.8 WAR. Given that there can only be 1000 WAR in a season, they have an error somewhere. Neil Weinberg suggested “the reason for this is because no one has gotten hurt yet, so the depth charts are oversampling PA/IP from good players”. If this is the case, then Fangraphs is being overly optimistic projecting health from established players. The Orioles are depending on a number of star players and have limited quality depth available, so they are more vulnerable to injuries than other clubs. Such a flaw would overrate the Orioles’ ability.

Fangraphs doesn’t think much of the Orioles starting pitching. It is ranked third worst in WAR and are projected to have the second lowest ERA. They also project the Orioles’ starters to throw the second lowest amount of innings in the majors, even fewer than teams like the Reds and Phillies. The Orioles rotation performed poorly in 2015, and lost Wei-Yin Chen in free agency so such an event would be unfortunate but also plausible.

Fangraphs is higher on the Orioles’ bullpen. It is ranked sixth in WAR, and 20th in ERA. The Orioles’ rotation isn’t expected to throw many innings and therefore the bullpen will have to take over the slack. As a result, the Orioles top relievers of Britton, O’Day, Brach and Givens are projected to throw only 230 of the 523 innings thrown by the bullpen. The more innings thrown by long relievers and non-elite arms, the more runs that the Orioles’ bullpen will allow to opposing teams. It’s very possible for the Orioles to have elite relievers in the bullpen and still have a poor bullpen ERA.

It makes sense to presume that the Orioles pitching will struggle due to the weakness of the Orioles rotation, but not that they would be the second worst in the majors especially given their decent bullpen. I decided to create a table showing the total runs allowed by each teams’ pitching, and the amount of earned runs allowed by starters and relievers. In addition, by subtracting total runs from earned runs, I was able to derive the total amount of unearned runs allowed by each team.

I found that the Orioles rotation was projected to allow roughly 455 earned runs or the fourth most in the majors. Their bullpen was projected to allow 218 earned runs or the seventh most in the majors and all told the Orioles were expected to allow 674 earned runs good for the fifth most in the majors. It wouldn’t be particularly surprising if the Orioles allowed that many earned runs. The Orioles allowed 642 in 2012, 678 in 2013, 558 in 2014 and 646 in 2015 so allowing 674 in 2016 would be within reason even if I’d expect 650. The chart looks like this.


The reason why the Orioles do so poorly is that they are projected to allow 86 unearned runs or tie for worst in the majors with the Blue Jays. This is surprising as the Orioles allowed 31 in 2013, 35 in 2014 and 47 in 2015. 86 unearned runs is more then the Orioles allowed in 2014 and 2015 combined. No team has allowed 86 unearned runs or more in the past three seasons and these projections project that each team will allow an average of 75 unearned runs this year versus roughly 50 from 2013-2015.

The numbers look even more bizarre when looking at the five teams projected to allow the most unearned runs in 2016. The Royals are actually tied for third with 84 despite having an elite defense and the other two clubs are the Diamondbacks and Reds. All five of these clubs are projected to have top eight defenses in 2016 as measured by Fangraphs Field Metric and would probably be expected to have some of the lowest unearned run totals in the majors.


Neil did argue that fielding isn’t equal to unearned runs. This is completely true, but the field metric does have a correlation with unearned runs. The correlation between the fielding metric and unearned runs was -.436 in 2013, -.491 in 2014, -.489 in 2015. This is the expected result as it suggests there's a moderate correlation between having good fielders and limiting unearned runs. One shouldn't expect a strong correlation because factors such as sequencing impact the amount of unearned runs a team allows and is independent of defense. Likewise, teams with good range will have good fielding scores but may have many errors.

For 2016, the projected correlation is +.487. This suggests that there’s the same amount of certainty between 2013-2015 and 2016, but Fangraphs is projecting that good fielding teams will allow more unearned runs. This is counter to historical data and basic common sense. Given the relationship between the field metric and unearned runs, it would seem likely that there is a faulty addition or subtraction operation in their projection model.

Using historical data and 2016 projected defense, I built two models that project how many unearned runs teams should be expected to allow in 2016. I found that teams like the Orioles and Royals should be expected to give up forty fewer unearned runs than currently projected while teams like the Pirates and Padres should probably allow roughly 60 unearned runs. If so, when taking into account the lower run environment, teams like the Royals should be expected to win an extra two games, while the Padres should be another two games worse. Two wins may appear to have only a minor impact, but the difference between the third-best team and third-worst team in the majors is only 17 wins and another six wins would change the Royals from being in the cellar to being in first place. Small changes have a larger impact then one may think.

The Fangraphs model has 1080 WAR instead of 1000 and projects each team will allow 75 unearned runs rather than only 50. They may want to consider looking into these problems and fixing them.

14 July 2015

Why You Can't Just Look at WAR to Determine a Player's Ability

The other day, I got into an argument about Rick Porcello. One person made that argument that if you believe in fWAR, Porcello has been good. He’s been worth 8.4 fWAR over the past 3.5 years or about roughly 2.4 fWAR per year, primarily due to a strong FIP and the ability to pitch a lot of innings. If one win costs $7.5 million then paying $20 million per year is a slight but not huge overpay. Writers at Fangraphs have also argued that Porcello is underrated, that he’s developed nicely into a 3-win player, that moving to Boston will make him better, that he deserved a huge payday, and that paying $20 million per year is reasonable. Paul Swydan, an author for Fangraphs, wrote an article in the Boston Globe suggesting that Porcello is the 13th-best pitcher in baseball.

On the other hand, I made the argument that Porcello is a slightly better version of Bud Norris. Let me explain why I made that argument and why just looking at WAR to decide pitchers' value isn’t always the best idea.

This first table compares Norris and Porcello’s performances from 2012-2015.


Porcello has a number of advantages. He’s been healthier so therefore he’s thrown more innings, but he also throws more innings per start. His win-loss record is slightly above .500 while Norris’s was slightly below .500 and they have roughly the same ERA. The main difference is that Porcello has a FIP that’s 0.4 runs lower than Norris and that’s why he has a considerably higher fWAR than Norris but a similar RA9_WAR.

This second table compares Norris and Porcello’s performance from 2012-2015 with the bases empty, runners on base, and runners in scoring position.


Porcello does a good job pitching with no one on base. He has a decent strikeout rate and more importantly an excellent walk rate. He gives up a standard home run rate, but it ends up resulting in fewer home runs than average due to a low fly ball rate. When no one is on base, Porcello is an ace. Meanwhile, Norris does a poor job in those situations. He gives up a lot of walks and has a horrific FIP of 4.68.

The problems start for Porcello when runners are on base. His K-BB% drops from 15% when the bases are empty, to 5.2% when a runner is on base, to 2.8% when a runner is in scoring position. The amount of fly balls that he gives up stays the same, but he also allows more home runs due to a higher HR/FB%. His HR/FB% is higher than average for reasons that will become clear later in the post. Unsurprisingly, his FIP goes from 3.25 with the bases empty, to 4.58 with a runner on base, to 4.8 with runners in scoring position.

Meanwhile, Norris improves when men are on base. His K-BB% jumps from 9.6% to 14.4% and his HR/9 rate drops from 1.3 with the bases empty, to 1 with a runner on base, to 0.85 with runners in scoring position. Unsurprisingly, Norris has a better FIP when pitching with men on base than when pitching without men on base.

The bottom line is that Norris becomes more effective when runners are on base while Porcello is less effective. The problem with that is that ERA measures what actually happens so by definition, it takes Porcello collapsing with runners on base into account. After all, that causes him to allow more runs which counts against his ERA. FIP doesn't have a way of differentiating between how Porcello does with men on base and with the bases empty. The formula presumes that a pitcher will perform the same with runners on base than with the bases empty and therefore doesn’t take into account the fact that Porcello does a terrible job pitching with the bases empty. It seems reasonable that this flaw means that in this case, ERA is a more effective estimator than FIP. At the very least, it indicates that FIP is a bad estimator to determine Porcello's performance. Honestly, if any of the two pitchers has had bad luck it’s probably Bud Norris, as one would expect him to have a lower ERA than his FIP which isn't the case.

Furthermore, Porcello’s performance in this regard has been reasonably consistent. This is what he’s done from 2012-2015.


His performance has been pretty much consistent. It's true that he did better with the bases empty in 2013 than he has in previous years. He has also performed slightly worse with the bases empty in 2015. Likewise, when men are on base the numbers are also reasonably consistent. His 2015 FIP is a bit worse due to an elevated HR/FB% but his 2015 xFIP is in line with normal figures.

The only case where there’s a significant change is in 2014 when runners are in scoring position. In those situations, his FIP was 3.9 while his average FIP from 2012-2015 was 4.8. But the reason why his FIP was so good in 2014 in those situations was because of a 4.8% HR/FB rate and not because he was able to fix his poor K-BB%. A 4.8% HR/FB rate is not sustainable and indeed his xFIP for 2014 with RISP is similar to his 2012-2015 average.

Basically, the data show that Porcello hasn’t had a good K-BB rate with men on base in any year from 2012 to 2015 and that his success in 2014 was due to avoiding home runs with men on base. That's not a strategy for success.

This next table is created with data from ESPN's Stats and Information portal and further shows how Porcello has done from 2012-2015 with men on base.


It tells pretty much the same story. I'm including it because it has statistics like OPS and wOBA that may be more useful to the user, It also shows how a deflated BABIP also contributed to Porcello’s success in 2014. Looking at Porcello’s performance in 2015, we can pretty safely say that his good fortune didn’t continue. A pitcher doesn't often have a .684 OPS with an 11.60 K% and a 9.00 BB%.

One might wonder why Porcello was able to give up fewer home runs in 2014 than he did in other seasons. This next table, using data provided by ESPN Stats and Information, shows how many fly balls Porcello allowed with RISP from 2012-2015.



Porcello's fly balls weren’t hit as hard in 2014 with RISP as they were in 2013 and 2015 but they were hit as hard as they were in 2012. That could be seen as a good sign, but the problem is that Porcello only allowed 40 in those situations in 2014. This is an awfully small sample and in light of his 2015 results, it appears that it was just fortunate chance. This is especially supported by the fact that his fly balls were hit roughly just as hard with men on base in 2014 as they were in 2012 and 2013. He's been pounded pretty badly in 2015.

The next question is why does Porcello struggle to get strikeouts when runners are in scoring position? This is easily answered by looking at the results of his pitches over the period using data from ESPN Stats and Information. Here’s a chart.


Porcello throws more strikes when the bases are empty than when there are runners in scoring position while also allowing fewer balls being put into play. This results in him having a higher percent of called strikes when the bases are empty than when runners are in scoring position as well as also allowing more foul balls. It would seem that batters are better able to predict where his pitches will go when batters are in RISP chances than not. All in all, more strikes and fewer balls put into play results in more strikeouts and fewer walks when no one is on base.

This next table shows how batters perform against Porcello’s pitches.


Porcello appears able to throw his fastball for strikes and can use it to get strikeouts. The problem is that batters absolutely annihilate them when they put them into play. Batters hit the pitch so hard in fact, that it probably is a bad idea to throw it. In addition, batters also crush his curve/slider when they put those pitches into play. Those pitches appear to be slightly successful when no one is on base but result in absolute disaster when men are on base. The bottom line is that he only has an effective sinker and changeup. It turns out that there's a major difference between being able to throw five pitches and being able to throw five pitches well.

Naturally, the Red Sox have adjusted to this fact by changing what pitches he throws. This next table shows the percentages of each pitch he throws each year.


For some reason, the Red Sox have decided that Porcello should throw his fastball more often and that he should throw his changeup and sinker less often. Or they’ve decided he should throw his worst pitch rather than his best pitches. I have no idea why they'd resort to this strategy but it turns out that having a pitcher throw his worst pitches more often results in him having a worse year than average.

This suggests that his results in 2015 have been earned and that they aren't representative of how he could perform used properly. It also makes one wonder whether what the Red Sox are planning and whether they can use him properly.

If one just looks at WAR, FI,P and health, then Porcello appears to be a good pitcher. He’d almost definitely be considered above average if not a solid No. 2. Given that he's been healthy, it would seem reasonable to give him a large contract based on his prior performance despite his poor ERA.

However, if one takes a more in-depth look at his stats, it quickly becomes clear that he’s terrible when men are on base or in scoring position and was successful in 2014 solely due to luck with home runs. It seems he doesn’t have a viable fastball, is unable to throw strikes in the clutch, and that his ERA is probably a better predictor of his true ability than his FIP. This probably means that he’s a No. 5 starter and his true ability is limited. I suppose he may be better than Bud Norris but certainly not worth $20 million per year.

That’s exactly why one can’t just look at WAR to gauge ability. While WAR is helpful, it’s solely a number that summarizes a pitcher's performance without providing much insight into why a pitcher performs the way he does. Sometimes that insight makes it clear that a pitcher isn’t as good as one would otherwise think.

11 September 2014

The Problem With WAR

Jeff Passan recently wrote an article about WAR that sparked some major debates on Twitter. In response, Dave Cameron wrote this article. Reading these articles got me thinking about WAR and one of the major issues with it.


When I was a kid I didn’t know about WAR, wRC+, ISO, wOBA or other advanced stats that quantify a players offensive contributions. I’m not even sure I was familiar with OBP, SLG or OPS. I judged a players’ offensive ability based on his batting average, home runs and RBIs. If I really wanted to look closely at a player I would maybe consider how many doubles and triples he hit as well as how many times he walked and struck out. 

By themselves these stats have limited utility but they are still useful. Batting average does give a basic idea of how often a player gets on base even if it's not as good as OBP. Home Runs and RBI do indicate how much power a player has even if they're not as useful as ISO or even SLG. Most people would have agreed back than that a home run is more valuable than a triple, a triple more valuable than a double, a double more valuable than a single and a single more valuable than a walk.

Due to the creation of these basic statistics one understands the necessity of comparing different facets of offense. For example, is a player batting .260/30/80 more valuable than a player batting .310/15/50? You can't really answer the question with those tools and in fact the answer doesn't really matter. What’s important is that these basic statistics allow us to ask the question. It becomes clear that the number of hits as well as the quality of hits matter.

The fact that these basic offensive statistics exist means that it’s easier for us to understand more advanced offensive statistics. Take SLG for example. It makes sense that certain hits are more valuable than others. And it makes logical sense that a home run that is worth four bases is four times the value of a single that is worth one base. Basic offensive statistics allow us to build that model. Once we start assigning arbitrary values to home runs, triples, doubles, singles, walks etc it becomes easier to understand how we can assign values to them based on historical data.

Furthermore, most basic offensive statistics are easy to define. There’s little difficulty in defining when a player hits a double. It’s reasonably straightforward because either the batter is on second or he isn’t. One can argue about the value of a double but in the vast majority of cases it’s hard to argue whether or not a double occurred.

Decades of statistics has accustomed us to quantifying the value of offensive production as well as give us easily defined and understood tools to do so.

But basic statistics aren’t nearly as helpful quantifying defensive contributions. The only basic defensive statistics are errors, putouts, assists and fielding percentage. Putouts and assists are fairly straightforward but errors are often subjective. Basic defensive statistics aren’t as objective as basic offensive statistics.

Furthermore, you can’t really use fielding percentage to compare two players playing the same position let alone compare players playing at different positions. Fielding percentage simply doesn't quantify range. It doesn’t quantify whether one fielder made a lot of excellent plays or a few excellent plays. I would say that it’s similar to batting average. Batting average is a helpful basic statistical stat but one wouldn’t use it by itself to quantify offensive contributions.

All I’m saying is that I’d feel far more comfortable claiming that a player with a .330/30/120 line is better offensively than a player with a .210/15/50 line than I would claiming a player with a .985 fielding percentage is better defensively than a player with a .975 fielding percentage.

Basic defensive statistics tell us very little about defense. It was pretty much impossible for the casual fan to objectively quantify defense prior to advanced defensive statistics for large populations of players. After all, most casual fans simply aren’t watching a majority of the games for a majority of teams. It would be pretty time consuming.

This means that advanced statistics like UZR were really the first attempts to actually objectively quantify defense. The problem is that basic defensive statistics don’t give people a frame of reference. It’s hard to explain why defense is as valuable as it is considered by UZR because it has never been quantified before. People watching games probably would agree that Cal Ripken is a good defender. No one watching games would have claimed that Cal Ripken’s glove is worth 15 runs a season and seriously meant exactly 15 runs. And without using an advanced statistic like UZR it would be hard to tell whether being worth 15 runs defensively is good or bad.

This makes explaining defensive statistics a challenge because people don't really have the background to understand them. The way to explain them is by understanding that the concept is difficult and by openly disclosing all of the data and the formulas. The concept behind UZR has been explained as comparing the play that actually happened (hit/out/error) to data on similarly hit balls in the past to determine how much better or worse the fielder did than the "average" player.

But the public isn’t informed of how likely it is for a fielder to successfully field a certain play. For a given player I can see whether he was successful at any given at bat. I can’t tell whether a player was successful defensively on any given play. If a ball is hit to the outfield I don't know how likely it is that a different fielder would have made the play according to UZR. Unlike with offense, there is no play index for UZR. It is impossible to determine a fielders’ UZR when at home or away. It is impossible to determine a fielders’ UZR for a given month. In short, UZR and most of the rest of advanced defensive metrics pretty much boil down to one number for a given player and you can either take it or leave it. They may be the best numbers that we have available. But we’re pretty much taking their creators word that they’re accurate. It shouldn’t come as a surprise that many people simply choose to leave it.

The problem with WAR is that it relies on defensive statistics that haven't been adequately explained. Unlike statistics like FIP and WAR it is impossible to derive UZR numbers on our own. It doesn't really matter whether UZR is accurate or not. The point is that we have no background to determine its accuracy and it is impossible to test it ourselves.
 
The discussion about WAR has nothing to do with whether position players deserve 57% of WAR or 52% of WAR. It isn't about whether pitchers are responsible for 93% of run preventation. It's about the fact that defensive metrics aren't adequately available to the public for study.

Without further disclosure people are going to resist accepting advanced defensive metrics and as a result are going to question WAR. If the public is unable to repeat the methodology then these metrics have similarities to opinion. People aren’t going to fully trust a metric that can’t be fully understood or duplicated and simply shouldn't be asked to do so.