Showing posts with label fWAR. Show all posts
Showing posts with label fWAR. Show all posts

20 April 2017

Why WAR-Based Systems Underestimate Elite Relievers

People using WAR based systems to value players typically think that relievers are overvalued by baseball clubs.  Way back in 2010, an article printed on Fangraphs made the following claim:

WAR, as you probably know, doesn’t think much of relief pitchers. The very best relievers in the game are generally worth +2 to +2.5 wins over a full season, or about the same as an average everyday player. This has caused quite a few people to state that WAR doesn’t work for relievers, because the results of the metric don’t match what they believe to be true about relief pitcher value. I think it works just fine.
While the quality of their work is very high, the quantity is low, which limits their total value. It’s nearly impossible to rack up huge win values while facing less than 300 batters per season. Yes, each of those batters faced are more critical to a win than a regular batter faced, but this is accounted for in WAR.

Since that article, there appears to be a consensus that the WAR based systems still seem to undervalue relievers compared to actual baseball clubs, but people are now trying to understand the reasoning for that gap. Ron Arthur of 538 noted that teams with a good bullpen or more likely to hold a one-run lead and therefore made a hypothesis that this may be a reason why relievers seem to be overvalued by front offices relative to the sabermetric consensus. An article in Baseball Prospectus asked why teams seem to be willing to spend millions on relievers when WAR(P) tells them that spending on relievers is a mistake. The author argues that WAR(P) doesn’t do a good job measuring reliever value and that Win Probability Added (WPA) helps explain teams reasoning. Finally, a more recent article in Fangraphs notes that there’s a sizable and growing gap between the public’s valuation of elite relief arms and the industry’s valuation.

I think the reason why WAR-based systems struggle to quantify the value of relievers is because of a focus on value over replacement rather than value over average. If one focuses solely on value over replacement, than a 1 WAR starter who pitches 200 innings is just as valuable as a 1 WAR reliever who pitches 60 innings. In this scenario, since a replacement player by definition produces 0 WAR, each player contributed 1 WAR more than a replacement player. And since a replacement level starter is better than a replacement level reliever, it makes sense to argue that the 1 WAR starter is even more valuable than the 1 WAR reliever.

However, the results are different if one focuses on value over average. Suppose there are two pitchers on two teams; one pitcher produces 2 WAR over 60 innings and another that produces 2 WAR over 200 innings. If we presume that each team’s pitchers throw 1450 innings total and that the average team produces 14.3 pitching WAR, then the team with the pitcher that threw 60 innings needs to earn 12.3 WAR over 1390 innings to be average while the other team needs to earn 12.3 WAR over 1250 innings to be average. This example illustrates how two pitchers can earn the same amount of WAR, but one pitcher can help his team more than the other pitcher.

Focusing on value over average is probably a better strategy than focusing on value over replacement. Each year, pitchers earn roughly 430 fWAR split out among all 30 teams. From 2000-2016, there has been only one team with negative pitching WAR. Ultimately, most teams have a large population of pitchers that are above replacement and therefore need to consider not only how best to maximize WAR but to minimize innings pitched. Innings are a finite constraint that aren’t given enough consideration in WAR-based systems.

Recently, maybe even this week, Fangraphs updated its methodology for how to determine pitching WAR. It’s complicated but a simple and incomplete version of their method is that they determine Wins Per Game Above Replacement (WPGAR), multiply this by the innings pitched and then throw in a few more adjustments. For our purposes, the only relevant fact is that their methodology focuses on replacement rather than average.

In order to determine a Wins Above Average metric, I determined that on average from 2010-2016, an average pitcher earns 1 WAR per 101 innings. Therefore, to determine Wins Above Average, I take a pitchers WAR and subtract from it (Innings/101). For a more complicated Wins Above Average metric, I’d use Fangraphs methodology for determining wins above average and see whether our numbers are similar, but I didn’t learn about their updated methodology in enough time to perform the necessary calculations.

This metric is friendlier towards elite relievers than Wins Above Replacement. For example, from 2010-2016, 2016 Zach Britton ranks 402nd in WAR but 184th in Wins Above Average.
There aren’t so many surprises in the top ten pitchers. The absolutely amazing Clayton Kershaw is ranked #1 despite throwing only 149 innings. Rich Hill ranks 10th with a strong 110 innings. But in the next top 10 pitchers, there are relief pitchers Jansen, Miller and Betances. Chapman is ranked 21st and Britton is ranked 25th.

I built a basic model converting Wins Above Average to projected salary to see how elite reliever salaries might compare with this metric instead of using Wins Above Replacement. The model could use some improvement and isn’t ready for primetime, but it seemed to indicate that relievers might be 30-40% more valuable using Wins Above Average instead of Wins Above Replacement. If so, this could help explain why elite relievers are receiving higher salaries than WAR-based systems suggest.

WAR-based systems presume that players replacing other players are only at replacement level. This presumption means that relievers that can earn 2 WAR while throwing in just sixty innings are valued equally to starters that earn 2 WAR if they throw in 200 innings. Until WAR-based systems can find a way to take production over a limited period into account, it seems plausible that they will continue to underestimate the value of elite relievers.

14 July 2015

Why You Can't Just Look at WAR to Determine a Player's Ability

The other day, I got into an argument about Rick Porcello. One person made that argument that if you believe in fWAR, Porcello has been good. He’s been worth 8.4 fWAR over the past 3.5 years or about roughly 2.4 fWAR per year, primarily due to a strong FIP and the ability to pitch a lot of innings. If one win costs $7.5 million then paying $20 million per year is a slight but not huge overpay. Writers at Fangraphs have also argued that Porcello is underrated, that he’s developed nicely into a 3-win player, that moving to Boston will make him better, that he deserved a huge payday, and that paying $20 million per year is reasonable. Paul Swydan, an author for Fangraphs, wrote an article in the Boston Globe suggesting that Porcello is the 13th-best pitcher in baseball.

On the other hand, I made the argument that Porcello is a slightly better version of Bud Norris. Let me explain why I made that argument and why just looking at WAR to decide pitchers' value isn’t always the best idea.

This first table compares Norris and Porcello’s performances from 2012-2015.


Porcello has a number of advantages. He’s been healthier so therefore he’s thrown more innings, but he also throws more innings per start. His win-loss record is slightly above .500 while Norris’s was slightly below .500 and they have roughly the same ERA. The main difference is that Porcello has a FIP that’s 0.4 runs lower than Norris and that’s why he has a considerably higher fWAR than Norris but a similar RA9_WAR.

This second table compares Norris and Porcello’s performance from 2012-2015 with the bases empty, runners on base, and runners in scoring position.


Porcello does a good job pitching with no one on base. He has a decent strikeout rate and more importantly an excellent walk rate. He gives up a standard home run rate, but it ends up resulting in fewer home runs than average due to a low fly ball rate. When no one is on base, Porcello is an ace. Meanwhile, Norris does a poor job in those situations. He gives up a lot of walks and has a horrific FIP of 4.68.

The problems start for Porcello when runners are on base. His K-BB% drops from 15% when the bases are empty, to 5.2% when a runner is on base, to 2.8% when a runner is in scoring position. The amount of fly balls that he gives up stays the same, but he also allows more home runs due to a higher HR/FB%. His HR/FB% is higher than average for reasons that will become clear later in the post. Unsurprisingly, his FIP goes from 3.25 with the bases empty, to 4.58 with a runner on base, to 4.8 with runners in scoring position.

Meanwhile, Norris improves when men are on base. His K-BB% jumps from 9.6% to 14.4% and his HR/9 rate drops from 1.3 with the bases empty, to 1 with a runner on base, to 0.85 with runners in scoring position. Unsurprisingly, Norris has a better FIP when pitching with men on base than when pitching without men on base.

The bottom line is that Norris becomes more effective when runners are on base while Porcello is less effective. The problem with that is that ERA measures what actually happens so by definition, it takes Porcello collapsing with runners on base into account. After all, that causes him to allow more runs which counts against his ERA. FIP doesn't have a way of differentiating between how Porcello does with men on base and with the bases empty. The formula presumes that a pitcher will perform the same with runners on base than with the bases empty and therefore doesn’t take into account the fact that Porcello does a terrible job pitching with the bases empty. It seems reasonable that this flaw means that in this case, ERA is a more effective estimator than FIP. At the very least, it indicates that FIP is a bad estimator to determine Porcello's performance. Honestly, if any of the two pitchers has had bad luck it’s probably Bud Norris, as one would expect him to have a lower ERA than his FIP which isn't the case.

Furthermore, Porcello’s performance in this regard has been reasonably consistent. This is what he’s done from 2012-2015.


His performance has been pretty much consistent. It's true that he did better with the bases empty in 2013 than he has in previous years. He has also performed slightly worse with the bases empty in 2015. Likewise, when men are on base the numbers are also reasonably consistent. His 2015 FIP is a bit worse due to an elevated HR/FB% but his 2015 xFIP is in line with normal figures.

The only case where there’s a significant change is in 2014 when runners are in scoring position. In those situations, his FIP was 3.9 while his average FIP from 2012-2015 was 4.8. But the reason why his FIP was so good in 2014 in those situations was because of a 4.8% HR/FB rate and not because he was able to fix his poor K-BB%. A 4.8% HR/FB rate is not sustainable and indeed his xFIP for 2014 with RISP is similar to his 2012-2015 average.

Basically, the data show that Porcello hasn’t had a good K-BB rate with men on base in any year from 2012 to 2015 and that his success in 2014 was due to avoiding home runs with men on base. That's not a strategy for success.

This next table is created with data from ESPN's Stats and Information portal and further shows how Porcello has done from 2012-2015 with men on base.


It tells pretty much the same story. I'm including it because it has statistics like OPS and wOBA that may be more useful to the user, It also shows how a deflated BABIP also contributed to Porcello’s success in 2014. Looking at Porcello’s performance in 2015, we can pretty safely say that his good fortune didn’t continue. A pitcher doesn't often have a .684 OPS with an 11.60 K% and a 9.00 BB%.

One might wonder why Porcello was able to give up fewer home runs in 2014 than he did in other seasons. This next table, using data provided by ESPN Stats and Information, shows how many fly balls Porcello allowed with RISP from 2012-2015.



Porcello's fly balls weren’t hit as hard in 2014 with RISP as they were in 2013 and 2015 but they were hit as hard as they were in 2012. That could be seen as a good sign, but the problem is that Porcello only allowed 40 in those situations in 2014. This is an awfully small sample and in light of his 2015 results, it appears that it was just fortunate chance. This is especially supported by the fact that his fly balls were hit roughly just as hard with men on base in 2014 as they were in 2012 and 2013. He's been pounded pretty badly in 2015.

The next question is why does Porcello struggle to get strikeouts when runners are in scoring position? This is easily answered by looking at the results of his pitches over the period using data from ESPN Stats and Information. Here’s a chart.


Porcello throws more strikes when the bases are empty than when there are runners in scoring position while also allowing fewer balls being put into play. This results in him having a higher percent of called strikes when the bases are empty than when runners are in scoring position as well as also allowing more foul balls. It would seem that batters are better able to predict where his pitches will go when batters are in RISP chances than not. All in all, more strikes and fewer balls put into play results in more strikeouts and fewer walks when no one is on base.

This next table shows how batters perform against Porcello’s pitches.


Porcello appears able to throw his fastball for strikes and can use it to get strikeouts. The problem is that batters absolutely annihilate them when they put them into play. Batters hit the pitch so hard in fact, that it probably is a bad idea to throw it. In addition, batters also crush his curve/slider when they put those pitches into play. Those pitches appear to be slightly successful when no one is on base but result in absolute disaster when men are on base. The bottom line is that he only has an effective sinker and changeup. It turns out that there's a major difference between being able to throw five pitches and being able to throw five pitches well.

Naturally, the Red Sox have adjusted to this fact by changing what pitches he throws. This next table shows the percentages of each pitch he throws each year.


For some reason, the Red Sox have decided that Porcello should throw his fastball more often and that he should throw his changeup and sinker less often. Or they’ve decided he should throw his worst pitch rather than his best pitches. I have no idea why they'd resort to this strategy but it turns out that having a pitcher throw his worst pitches more often results in him having a worse year than average.

This suggests that his results in 2015 have been earned and that they aren't representative of how he could perform used properly. It also makes one wonder whether what the Red Sox are planning and whether they can use him properly.

If one just looks at WAR, FI,P and health, then Porcello appears to be a good pitcher. He’d almost definitely be considered above average if not a solid No. 2. Given that he's been healthy, it would seem reasonable to give him a large contract based on his prior performance despite his poor ERA.

However, if one takes a more in-depth look at his stats, it quickly becomes clear that he’s terrible when men are on base or in scoring position and was successful in 2014 solely due to luck with home runs. It seems he doesn’t have a viable fastball, is unable to throw strikes in the clutch, and that his ERA is probably a better predictor of his true ability than his FIP. This probably means that he’s a No. 5 starter and his true ability is limited. I suppose he may be better than Bud Norris but certainly not worth $20 million per year.

That’s exactly why one can’t just look at WAR to gauge ability. While WAR is helpful, it’s solely a number that summarizes a pitcher's performance without providing much insight into why a pitcher performs the way he does. Sometimes that insight makes it clear that a pitcher isn’t as good as one would otherwise think.

21 November 2014

Can We Trust Lough's Defense?


Before David Lough went on his hot offensive streak in the middle of last season, Jon made the argument that Lough is actually a starter in disguise due to his excellent defense. He noted that Lough’s defense has been consistent for each season and projects to make him nearly a 2 WAR player over a full season. Pat suggested yesterday that an outfield consisting primarily of Adam Jones, Steve Pearce, Alejandro De Aza, and Lough could be acceptable. Certainly understanding Lough's value would be helpful. The main question is whether or not one can believe that Lough's defense is as valuable as UZR suggests.

Jon also noted in his earlier article that Lough’s defense may have been consistent for each season but that he’s played a limited number of innings in the field. So I looked at all outfielders that played at least 500 innings in the outfield from 2012-2014, and saw some surprising players in the top 20 in total UZR. For example, Lough’s total UZR was good for 11th out of 193 outfielders despite ranking 117th in innings. Lorenzo Cain and Juan Lagares were in the top five in total UZR despite ranking 58th and 87th in innings. Craig Gentry ranked seventh despite playing 12 more innings than Lagares. Dyson ranked ninth playing only 113 more innings than Lagares. This trend continues for outfielders that played 500 or more innings in the outfield from 2010-2012. Gardner ranked #1 in UZR and 59th out of 183 in innings. Peter Bourjos ranked #2 in UZR and 70th in innings. If outfielders with a minimal amount of innings playing in the outfield (i.e small sample sizes) have some of the highest UZRs then is there a correlation between innings played on defense and UZR? If not, then it may not make sense to say that an outfielders defensive value will improve if he plays more innings.

In the same vein, it seems reasonable to presume that players with more plate appearances produce more runs than those with fewer plate appearances because teams are far more likely to demote a player that is struggling offensively than one that is not struggling offensively. In other words, there should be correlation between plate appearances and total offensive production.

It's possible to answer this question. For each player from 2002-2014, I downloaded the number of innings they played at each position, their UZR at each position, their total number of plate appearances, and their “batting” score as derived by Fangraphs. Fangraphs "batting" score attempts to quantify the value of an offensive player by determine the number of runs above average his offensive production was worth in a given season. Then, I did a correlation analysis between innings and UZR as well as between plate appearances and “batting”. If there is a relationship between innings and UZR then I would expect to see a moderate correlation between these variables and a similar correlation between those statistics and the correlation between plate appearance and “batting.” Here's a table with the results.


The results suggest that the average correlation between innings and UZR is about .077. Defense intensive positions such as second base, third base, shortstop, and center field have larger albeit still low correlations than those for positions like first base, left field, and right field. The results for the correlation between PAs and "batting" tell a different story.



The average correlation between PAs and “batting” was .376. The largest correlations were at offensive positions such as first base, left field, and right field while there was practically no correlation at shortstop. It should come as no surprise that teams don't give shortstops playing time based on their ability to hit.

These results suggest that there is correlation between offensive production and plate appearances but not between UZR and innings played in the field and therefore it doesn’t make sense to give a player credit for a high defensive score that occurs in only part of a season. This explains why outfielders  that only play a limited number of innings consistently have some of the highest UZRs.

There is another possible way to interpret these results. These results could mean that managers don’t value UZR highly while valuing offensive production. If this is the case then there should be a similar correlation between innings and the absolute value of UZR as well as plate appearances and the absolute value of the “batting” statistic.

The average correlation between innings and the absolute value of UZR is .65 while the average correlation between plate appearances and the absolute value of “batting” is .52. This suggests that there is a moderate to high correlation between innings and UZR. It’s just that players that play a lot of innings at a position have a high absolute value that is either positive or negative. In other words, managers care less about whether their players have a low UZR than they do about whether their player is producing offensively. If UZR does measure defense accurately then this suggests that managers don’t value defense as highly as offense especially at offense-oriented positions.

This test is inconclusive in determining whether one should feel comfortable projecting Lough to be an excellent defender based on his UZR and performance in a limited sample. This should make us pause before overly valuing a limited period of good defense. Indeed, players like Nyjer Morgan, Ryan Sweeney, Andrew Torres, Tony Gwynn Jr., Gerardo Parra, and Ben Revere are all guys that put up strong defensive numbers in a limited amount of innings and then came crashing back to earth, and Lough could follow the same path.

This suggests one of two things. Either teams have better defensive metrics than UZR and therefore don't give its ratings high credence or that teams simply value offense more than defense especially for corner outfielders. This second scenario is probably bad news for outfielders like Jason Heyward. If UZR isn't a good defensive metric then it's likely that Lough's value is highly exaggerated. If UZR is a good defensive metric then finding a corner outfielder with good defense is an easy and undervalued way to improve. If so then even if the Orioles don't trust Lough's defense they should consider finding a proven player with similar strengths.

Photo via Keith Allison

11 September 2014

The Problem With WAR

Jeff Passan recently wrote an article about WAR that sparked some major debates on Twitter. In response, Dave Cameron wrote this article. Reading these articles got me thinking about WAR and one of the major issues with it.


When I was a kid I didn’t know about WAR, wRC+, ISO, wOBA or other advanced stats that quantify a players offensive contributions. I’m not even sure I was familiar with OBP, SLG or OPS. I judged a players’ offensive ability based on his batting average, home runs and RBIs. If I really wanted to look closely at a player I would maybe consider how many doubles and triples he hit as well as how many times he walked and struck out. 

By themselves these stats have limited utility but they are still useful. Batting average does give a basic idea of how often a player gets on base even if it's not as good as OBP. Home Runs and RBI do indicate how much power a player has even if they're not as useful as ISO or even SLG. Most people would have agreed back than that a home run is more valuable than a triple, a triple more valuable than a double, a double more valuable than a single and a single more valuable than a walk.

Due to the creation of these basic statistics one understands the necessity of comparing different facets of offense. For example, is a player batting .260/30/80 more valuable than a player batting .310/15/50? You can't really answer the question with those tools and in fact the answer doesn't really matter. What’s important is that these basic statistics allow us to ask the question. It becomes clear that the number of hits as well as the quality of hits matter.

The fact that these basic offensive statistics exist means that it’s easier for us to understand more advanced offensive statistics. Take SLG for example. It makes sense that certain hits are more valuable than others. And it makes logical sense that a home run that is worth four bases is four times the value of a single that is worth one base. Basic offensive statistics allow us to build that model. Once we start assigning arbitrary values to home runs, triples, doubles, singles, walks etc it becomes easier to understand how we can assign values to them based on historical data.

Furthermore, most basic offensive statistics are easy to define. There’s little difficulty in defining when a player hits a double. It’s reasonably straightforward because either the batter is on second or he isn’t. One can argue about the value of a double but in the vast majority of cases it’s hard to argue whether or not a double occurred.

Decades of statistics has accustomed us to quantifying the value of offensive production as well as give us easily defined and understood tools to do so.

But basic statistics aren’t nearly as helpful quantifying defensive contributions. The only basic defensive statistics are errors, putouts, assists and fielding percentage. Putouts and assists are fairly straightforward but errors are often subjective. Basic defensive statistics aren’t as objective as basic offensive statistics.

Furthermore, you can’t really use fielding percentage to compare two players playing the same position let alone compare players playing at different positions. Fielding percentage simply doesn't quantify range. It doesn’t quantify whether one fielder made a lot of excellent plays or a few excellent plays. I would say that it’s similar to batting average. Batting average is a helpful basic statistical stat but one wouldn’t use it by itself to quantify offensive contributions.

All I’m saying is that I’d feel far more comfortable claiming that a player with a .330/30/120 line is better offensively than a player with a .210/15/50 line than I would claiming a player with a .985 fielding percentage is better defensively than a player with a .975 fielding percentage.

Basic defensive statistics tell us very little about defense. It was pretty much impossible for the casual fan to objectively quantify defense prior to advanced defensive statistics for large populations of players. After all, most casual fans simply aren’t watching a majority of the games for a majority of teams. It would be pretty time consuming.

This means that advanced statistics like UZR were really the first attempts to actually objectively quantify defense. The problem is that basic defensive statistics don’t give people a frame of reference. It’s hard to explain why defense is as valuable as it is considered by UZR because it has never been quantified before. People watching games probably would agree that Cal Ripken is a good defender. No one watching games would have claimed that Cal Ripken’s glove is worth 15 runs a season and seriously meant exactly 15 runs. And without using an advanced statistic like UZR it would be hard to tell whether being worth 15 runs defensively is good or bad.

This makes explaining defensive statistics a challenge because people don't really have the background to understand them. The way to explain them is by understanding that the concept is difficult and by openly disclosing all of the data and the formulas. The concept behind UZR has been explained as comparing the play that actually happened (hit/out/error) to data on similarly hit balls in the past to determine how much better or worse the fielder did than the "average" player.

But the public isn’t informed of how likely it is for a fielder to successfully field a certain play. For a given player I can see whether he was successful at any given at bat. I can’t tell whether a player was successful defensively on any given play. If a ball is hit to the outfield I don't know how likely it is that a different fielder would have made the play according to UZR. Unlike with offense, there is no play index for UZR. It is impossible to determine a fielders’ UZR when at home or away. It is impossible to determine a fielders’ UZR for a given month. In short, UZR and most of the rest of advanced defensive metrics pretty much boil down to one number for a given player and you can either take it or leave it. They may be the best numbers that we have available. But we’re pretty much taking their creators word that they’re accurate. It shouldn’t come as a surprise that many people simply choose to leave it.

The problem with WAR is that it relies on defensive statistics that haven't been adequately explained. Unlike statistics like FIP and WAR it is impossible to derive UZR numbers on our own. It doesn't really matter whether UZR is accurate or not. The point is that we have no background to determine its accuracy and it is impossible to test it ourselves.
 
The discussion about WAR has nothing to do with whether position players deserve 57% of WAR or 52% of WAR. It isn't about whether pitchers are responsible for 93% of run preventation. It's about the fact that defensive metrics aren't adequately available to the public for study.

Without further disclosure people are going to resist accepting advanced defensive metrics and as a result are going to question WAR. If the public is unable to repeat the methodology then these metrics have similarities to opinion. People aren’t going to fully trust a metric that can’t be fully understood or duplicated and simply shouldn't be asked to do so.

03 April 2014

Is Defense Overrated?



Fangraphs recently released some Inside Edge fielding data to the public. As stated in the article, Inside Edge scouts watch every play and determine how difficult it is to field on the following scale:

  • Impossible (0%)
  • Remote (1-10%)
  • Unlikely (10-40%)
  • About Even (40-60%)
  • Likely (60-90%)
  • Almost Certain / Certain (90-100%)

Unlike zone-based applications like DRS and UZR, Inside Edge defensive system results are determined primarily by scouts. This data gives insight to how scouts grade defense and determine similarities and differences between zone-based and scout-based systems.  

The Inside Edge fielding data available to the public states how many chances and the success rate for each fielder in every category listed above. For example in 2013, Nick Markakis had 112 impossible chances of which he converted zero, five remote chances of which he converted zero, five unlikely chances of which he converted two, six about even chances of which he converted all six, eighteen likely chances of which he converted fifteen and 289 almost certain chances of which he converted 288.
I used this data to create a metric I think of as Catches over Average (COA). It is derived by downloading all Inside Edge fielding data for each position and determining the average success rate for each of the six categories described above. Next subtract the average conversion rates for each category from the players performance and multiply by the number of chances. Finally, sum the results for each category. Many concepts incorporated in this metric were originally discussed here.

This metric has a number of shortcomings. It presumes that each chance in a category is of equal difficulty. It doesn’t measure how successful an infielder is at turning a double play. It doesn’t measure how successful an outfielder is at ensuring singles don’t turn into doubles or outfielder arm strength. Nor does it measure how well a catcher frames pitches or prevents passed balls or throws out potential base stealers. Inside Edge collects this data but doesn’t share it with the public. Despite these shortcomings this metric still can be used to compare Inside Edge fielding data results to UZR and DRS.

This table shows how some of the Orioles main players performed defensively in 2012 and 2013.


There are many similarities between COA and UZR or DRS. Matt Wieters, Manny Machado and JJ Hardy are considered excellent defenders while Ryan Flaherty and Chris Davis are considered above average. Each indicates that Mark Reynolds and Wilson Betemit are unable to play third base very well while Adam Jones is a poor defender in center field. The main difference is that Nick Markakis is a good defender according to COA and a bad one according to UZR or DRS.

In 2012, our outfield defense ranked 16th in the majors with a -1.73 COA while our infield defense ranked 7th with a 9.13 COA. In 2013, our outfield defense ranked 26th in the majors with a -2.40 COA while our infield defense ranked 2nd with a 30.9 COA. Having a full season of Manny Machado at third base instead of using Wilson Betemit and Mark Reynolds had a significant impact on our infield defense. This statistic indicates that the Orioles have excellent infield defense and mediocre outfield defense which is the common consensus.

The difference between the best and worst infield defense was 48 COA in 2012 and 52 COA in 2013 while the difference between the best and worst outfield defense was 30.5 COA in 2012 and 20 COA in 2013. These numbers seem low when compared to UZR especially when one considers that a catch doesn’t equal a run. In order to do a full comparison it is necessary to determine how many catches equal a run.

Tom Tango states that each catch is worth .8 runs. This implies that the amount of runs saved can be determined by multiplying the value of a catch by a players COA. I’ll refer to this stat as Defensive Runs over Average or DROA. *

 It was brought to my attention that Tom Tango has discussed the value of saving a play. Originally, I quoted Michael Lichtman who stated that a typical outfield hit is worth .83 runs and simply estimated the value of an infield hit. I believe that estimate was incorrect.  

It is possible to compare DROA to UZR. The table below shows the ranges for UZR and COA for each position.

Position COA Range Value Of Catch DROA Range UZR Range UZR/DROA Range
1B 16.56 0.80 13.25 30.90 2.33
2B 25.94 0.80 20.75 31.70 1.52
3B 26.01 0.80 20.80 48.00 2.31
SS 25.22 0.80 20.18 45.20 2.24
CF 14.60 0.80 11.68 42.50 3.64
LF 12.15 0.80 9.72 31.30 3.22
RF 13.31 0.80 10.65 47.10 4.42

The range for UZR is about 2.25 times larger than DROA for infielders (except second base) and 3.75 times larger than DROA for outfielders. Instead of a top defender like Machado being worth 3 wins defensively this stat indicates that he’s worth one win. This implies that defense is less valuable than one may have thought based on UZR. **

** Due to the new data, I've updated the ranges quoted above.

The good folks at Retrosheet attempt to document every baseball game played. They share their data with the public provided that the following disclaimer is used:

Information used here was obtained free of charge from and is copyrighted by Retrosheet. Interested parties may contact Retrosheet at www.retrosheet.org

I downloaded the data from the 2012 and 2013 seasons from their site and determined how many balls in play allowed by each team. I split the balls in play into outs and non-outs (including hits and bases reached on error). Then I did the same thing with the Inside Edge Data. Below are the results.


The numbers of outs for balls in play are different in each dataset by a minimal amount while there are more than twice as many balls in play in the Retrosheet dataset for each team than in the Inside Edge dataset. The average team had a BABIP of .298 using the Retrosheet data and a .188 using the Inside Edge data. The Inside Edge data is omitting a large number of hits.

UZR and DRS split up the whole field into zones. Each ball in play has to fall into a zone which is the responsibility of at least one but potentially more fielders. If a ball isn’t caught then it is someone’s fault and it impacts their UZR/DRS rating. The Inside Edge data seems to indicate that part of the field is no fielders’ responsibility. Other parts of the field may be a fielders’ responsibility but an uncaught ball hit there has no impact on his defensive rating if it’s considered an impossible chance like most non-out balls in play. Balls hit into play that do have an impact of a fielders rating are usually ones that every fielder can field with nearly perfect accuracy. As a result, there is less variation in defensive ratings and an elite player according to UZR may be worth 3 wins defensively while an elite player according to DROA is worth 1 win defensively.

It is possible to determine whether UZR or COA is more likely to predict future performance by doing a correlation analysis. I found that UZR had a correlation of .393 from 2012 and 2013 while COA had a correlation of only .249. This suggests that UZR is more likely to predict future performance though neither is particularly accurate. I also found that while range runs above average is the best factor to predict UZR that error runs above average is the best factor to predict COA which supports the above paragraph. This explains why players similar to Nick Markakis that have limited range but make few errors have much higher defensive ratings using Inside Edge data rather than UZR while players like David Lough that have excellent range but make errors have much higher defensive ratings based on UZR rather than Inside Edge.

The average AL team allowed 696 runs in 2013. According to UZR, the Royals defense (best in the AL) was worth 68 runs. This means that a team with average pitchers and the Royals defense would have allowed the fourth fewest runs in the American League. Using DROA, the Orioles defense (best in the AL) was worth 16 runs. A team with average pitchers and the Orioles defense would have given up the ninth largest amount of runs or still been about average. The amount of importance that UZR places on defense as opposed to Inside Edge is huge.

If UZR is correct in its rating of defense then focusing on defense should be a cost-effective way to improve performance. If Inside Edge is correct then excellent defense players offer smaller contributions than UZR and WAR would suggest. I believe that the Inside Edge data suggests that defense is overrated.

25 September 2012

Was Promoting Manny Machado the Right Move for 2012?

Despite his top prospect status, I don't think the expectations for what Manny Machado would do in an Orioles' uniform this season were all that high. With Wilson Betemit on the DL though, the team need a third-baseman and decided to call up the 20 year-old from Double-A to fill in.

Machado started out well, hitting 3 home runs in his first 4 games, but since then has hit just .254/.267/.345 in 147 plate appearances (36 games). A 31 to 3 strike-out to walk ratio (overall) just isn't going to get it done in the Majors, but you don't need to be a plus hitter when you can handle the glove-work. As a shortstop, Machado unsurprisingly has solidified the hot corner for the O's - he's the only player with a positive UZR at the position this year (+3 runs), and is a big improvement over the likes of Betemit (-6 UZR) or Mark Reynolds (-5 UZR).

And even though the offensive production isn't there, given the alternatives it's still decent. Machado's .302 wOBA has translated into -2.2 runs relative to average in his 163 PA. Without Manny, the O's would have potentially had to go with a combination of Robert Andino, Omar Quintanilla, and Ryan Flaherty at third. The way they've hit this year, those three guys are at around -7, -3, and -6 runs in 163 PA, respectively. Combine that with losing some bench flexibility and potential platoon opportunities, and it seems fair to say that Machado has provided the team with upwards of 5 runs with the bat over what they would have otherwise been getting. With how close the AL East (and Wild Card) might be, that is certainly relevant.

Even if you ignore his hot start and assume he hit .254/.267/.345 all year, he'd still be a little above replacement level as a player; he's at +0.7 fWAR with -2.2 batting runs, so that translates into +0.2 fWAR with -7 batting runs with that triple slash. And that -7 batting runs is about what Andino would be expected to do over the same time period, so even the "slumping" Manny Machado doesn't really cost the Orioles anything.

More generally, this is probably good experience for Machado. He's holding his own in the big leagues despite not even getting a full year at Double-A. The lack of walks is a problem, but he hasn't been a complete hacker - he's swinging at pitches out of the strike-zone only slightly more often than league average (according to FanGraphs). Pitchers are pounding the zone against him though, and he's swinging and missing a fair bit. Still, there's enough positive signs to feel good about what Machado will be able to do in the near future - and, though scouting isn't my thing, he sure looks good out there.