Showing posts with label relievers. Show all posts
Showing posts with label relievers. Show all posts

20 April 2017

Why WAR-Based Systems Underestimate Elite Relievers

People using WAR based systems to value players typically think that relievers are overvalued by baseball clubs.  Way back in 2010, an article printed on Fangraphs made the following claim:

WAR, as you probably know, doesn’t think much of relief pitchers. The very best relievers in the game are generally worth +2 to +2.5 wins over a full season, or about the same as an average everyday player. This has caused quite a few people to state that WAR doesn’t work for relievers, because the results of the metric don’t match what they believe to be true about relief pitcher value. I think it works just fine.
While the quality of their work is very high, the quantity is low, which limits their total value. It’s nearly impossible to rack up huge win values while facing less than 300 batters per season. Yes, each of those batters faced are more critical to a win than a regular batter faced, but this is accounted for in WAR.

Since that article, there appears to be a consensus that the WAR based systems still seem to undervalue relievers compared to actual baseball clubs, but people are now trying to understand the reasoning for that gap. Ron Arthur of 538 noted that teams with a good bullpen or more likely to hold a one-run lead and therefore made a hypothesis that this may be a reason why relievers seem to be overvalued by front offices relative to the sabermetric consensus. An article in Baseball Prospectus asked why teams seem to be willing to spend millions on relievers when WAR(P) tells them that spending on relievers is a mistake. The author argues that WAR(P) doesn’t do a good job measuring reliever value and that Win Probability Added (WPA) helps explain teams reasoning. Finally, a more recent article in Fangraphs notes that there’s a sizable and growing gap between the public’s valuation of elite relief arms and the industry’s valuation.

I think the reason why WAR-based systems struggle to quantify the value of relievers is because of a focus on value over replacement rather than value over average. If one focuses solely on value over replacement, than a 1 WAR starter who pitches 200 innings is just as valuable as a 1 WAR reliever who pitches 60 innings. In this scenario, since a replacement player by definition produces 0 WAR, each player contributed 1 WAR more than a replacement player. And since a replacement level starter is better than a replacement level reliever, it makes sense to argue that the 1 WAR starter is even more valuable than the 1 WAR reliever.

However, the results are different if one focuses on value over average. Suppose there are two pitchers on two teams; one pitcher produces 2 WAR over 60 innings and another that produces 2 WAR over 200 innings. If we presume that each team’s pitchers throw 1450 innings total and that the average team produces 14.3 pitching WAR, then the team with the pitcher that threw 60 innings needs to earn 12.3 WAR over 1390 innings to be average while the other team needs to earn 12.3 WAR over 1250 innings to be average. This example illustrates how two pitchers can earn the same amount of WAR, but one pitcher can help his team more than the other pitcher.

Focusing on value over average is probably a better strategy than focusing on value over replacement. Each year, pitchers earn roughly 430 fWAR split out among all 30 teams. From 2000-2016, there has been only one team with negative pitching WAR. Ultimately, most teams have a large population of pitchers that are above replacement and therefore need to consider not only how best to maximize WAR but to minimize innings pitched. Innings are a finite constraint that aren’t given enough consideration in WAR-based systems.

Recently, maybe even this week, Fangraphs updated its methodology for how to determine pitching WAR. It’s complicated but a simple and incomplete version of their method is that they determine Wins Per Game Above Replacement (WPGAR), multiply this by the innings pitched and then throw in a few more adjustments. For our purposes, the only relevant fact is that their methodology focuses on replacement rather than average.

In order to determine a Wins Above Average metric, I determined that on average from 2010-2016, an average pitcher earns 1 WAR per 101 innings. Therefore, to determine Wins Above Average, I take a pitchers WAR and subtract from it (Innings/101). For a more complicated Wins Above Average metric, I’d use Fangraphs methodology for determining wins above average and see whether our numbers are similar, but I didn’t learn about their updated methodology in enough time to perform the necessary calculations.

This metric is friendlier towards elite relievers than Wins Above Replacement. For example, from 2010-2016, 2016 Zach Britton ranks 402nd in WAR but 184th in Wins Above Average.
There aren’t so many surprises in the top ten pitchers. The absolutely amazing Clayton Kershaw is ranked #1 despite throwing only 149 innings. Rich Hill ranks 10th with a strong 110 innings. But in the next top 10 pitchers, there are relief pitchers Jansen, Miller and Betances. Chapman is ranked 21st and Britton is ranked 25th.

I built a basic model converting Wins Above Average to projected salary to see how elite reliever salaries might compare with this metric instead of using Wins Above Replacement. The model could use some improvement and isn’t ready for primetime, but it seemed to indicate that relievers might be 30-40% more valuable using Wins Above Average instead of Wins Above Replacement. If so, this could help explain why elite relievers are receiving higher salaries than WAR-based systems suggest.

WAR-based systems presume that players replacing other players are only at replacement level. This presumption means that relievers that can earn 2 WAR while throwing in just sixty innings are valued equally to starters that earn 2 WAR if they throw in 200 innings. Until WAR-based systems can find a way to take production over a limited period into account, it seems plausible that they will continue to underestimate the value of elite relievers.

19 August 2016

Zach Britton is the 2nd Best Pitcher in the American League

Begrudgingly, nearly everyone in baseball data science will say that they do not completely comprehend the value of relief pitching.  Often, you will hear how dominant relievers are incapable of starting.  That is pretty much true.  Often, you will hear that closers are not employed at the most consequential moments and often are given a clean slate when entering a game.  That is also true.  Often, you will hear how utterly confused analysts are when a reliever is paid big money, given a lot of years, or is acquired in exchange for multiple, notable prospects.  Often, you will hear how several teams put together dominant bullpens that are collections of spare, rubbish heap arms, which is also true.

If you simply read what most data analysts say, relief pitching continues to be overrated and dominant relievers are failed starters.  Of course, this statement seems a little silly, right?  How many failed starters are there each year at the MLB level?  At least 30.  Very few of them become dominant relievers let alone dependable middle relief arms.  There certainly is something more to it.  Relief arms tend to need one or two above average pitches and at least mediocre control.  The reality is that this is not a common thing among those failed starters.

Furthermore, maybe our ability to appreciate relievers with metrics like WAR simply measures their value inadequately.  In this post, I try to look at things from a couple different angles.  First, a somewhat traditional way looking at saves and blown saves.  To do this, I took the players with at least 60 save opportunities from 2013 to 2015 and selected the top 20 players.  I assumed replacement level closing was the average success rate of the bottom 5 of those top 20 closers.  This might actually be an overestimation simply because to rack up 60 save opportunities, you have to be an arm that a club has been devoted to.

What we find is that the bottom rung long term closer blows 6.5 games per 40 opportunities (84% success) every year.  A top five closer blows 3.1 games per 40 opportunities (92% success) every year.  That is a difference of 3.4 games.  A blow game does not mean a loss.  I would guess that a blown save in a closing situation may be a loss 70% of the time, which would mean that the bottom run closer would lose 2.4 more games per year.  At a cost per win of 7 MM, that would suggest on average that the elite closing arm is worth about 18 MM, which suggests that maybe closers are somewhat undervalued in the market.

Zach Britton is currently 37 for 37.  Based on the above, we would expect a bottom run closer to blow 5.92 games over that stretch.  If we depreciate those blown saves in a conversion over to losses, then we have 4.1 losses.  This would suggest that Britton has actually been worth over 4 wins so far this season.  This would nestle him right behind Corey Kluber's 4.3 fWAR for second in the AL among all pitchers.

Maybe the simplicity of looking only at save opportunities leaves your brain unmoved and your heart cold.  Well, we can dive into RE/24.  RE/24 looks at run expectancy before and after each event in a game and attributes those to a pitcher without consideration of anything other than run expectancy.  A starting pitcher can benefit simply by racking up successful innings and a reliever benefits by coming in to high leverage situations.  RE/24 often is most unfair to closers who tend to enter the game with a clean slate and somewhat indulgent to middle relievers who successfully enter games with men on base.

Anyway, I batched all starters together and all relievers together for each team, ran those variables along with RE/24 for team batting, and regressed all of that against team wins.  What I found is that relief RE/24 was 78% the value of starting pitcher RE/24, which is not accounted for with innings pitched.  I then took the RE/24 AL pitcher leaderboard and scaled down relief pitcher RE/24.  Next, I accounted for park factors in home and away stadiums as well as the defensive ability for each team.  What resulted was a RE/24(x) metric that I created.  Here is that leaderboard:

RE24(x)
1
27.0
2
25.3
3
22.8
4
22.6
5
19.4
6
18.5
7
18.3
8
16.8
9
16.4
10
15.7
11
15.3
12
15.2

Britton shows up as sixth on this board.  He is not exactly challenging the leaders much, but he still shows he is in the conversation for Cy Young.

Of course, the argument might wind up being that while Britton excels at closing, so would several of the other pitchers on this board.  One way to look at that would be to see what exactly the impact of higher velocity might have on a pitcher's success.  An increase of 1 mph in general decrease a player's FIP by about 0.40.  Not all starters when pressed into relief roles enjoy an increase in velocity, but lets be kind and simply assume all starters would see a jump of 3 mph that suggests a FIP improvement of 1.20.

Our leaderboard using the players above would yield:

xrFIP
1 Corey Kluber 1.81
2 Zach Britton 2.00
3 Aaron Sanchez 2.09
4 Danny Duffy 2.13
5 Jose Quintana 2.22
6 Michael Fulmer 2.26
7 Chris Sale 2.29
8 Cole Hamels 2.46
9 Justin Verlander 2.46
10 J.A. Happ 2.69
11 Marco Estrada 2.96
12 Chris Tillman 3.07
All of this tends to suggest that maybe all those who are upset with Zach Britton being in the Cy Young conversation might be relying a bit too heavily on unsteady data science analysis with respect to relief pitching.  Many analysts may be harping loudly about concepts and ideas that have been firmly in place within the sabermetric community for over a decade and may be forgetting that this is an area of baseball that has yet to firmly establish what exactly the value is of a closer in a definitive way.

Perhaps in the coming years, data science will find ways to measure relief pitching quality that has a more substantial methodology than we currently have.  Maybe that methodology will show that closing is not the realm of inadequate pitchers, but more perhaps something completely different.  The skill set needed to be an elite pitcher differs from those who start and maybe that means these are truly two different positions within the pitcher class instead of a first and second tier.

Or maybe not.

12 December 2014

Don't Trust a Reliever Farther Than You Can Throw Him

Relievers have been in high demand this offseason. Potential closer candidate Andrew Miller received a four-year deal for $36 million while David Robertson received four years and $46 million. Non-closers like Pat Neshek received two years and $12.5 million while Luke Gregerson got three years and $18.5 million. The value of a good bullpen was proven when the Royals rode their top guys to the World Series last season. But many believe that even the best relievers can be highly volatile and therefore teams should be skeptical of offering them large contracts. In order to see whether this is valid, I looked at all of the 301 relievers from 2010-2012 who threw at least 50 innings in relief during that stretch, and compared their performance for 2013 and 2014 to all 265 relievers who threw at least 40 innings in relief using ERA, WAR, and RA9_WAR.

The chart below shows how each of the top 50 relievers according to each of the statistics from 2010 to 2012 performed in 2013 and 2014. 


Only 31.5% of the top 20 relievers and 24.1% of relievers ranked 21-50 according to ERA in 2010-2012 were top 50 relievers in 2013 and 2014. The average top 50 reliever ended up being an asset to the bullpen but wasn’t the star reliever that the team signing him was hoping to receive. The average ERA of roughly 3.15 looks like it's acceptable. However, 100 out of the 265 relievers who threw at least 40 innings in relief from 2013 to 2014 had an ERA under 3.00. An ERA of 3.15 is slightly better than the median ERA of 3.35 for all qualified relievers.

45% of the top 20 relievers and 17.2% of relievers ranked 21-50 according to WAR in 2010-2012 were top 50 relievers in 2013 and 2014. The value of their contributions decreased dramatically from 2010-2012 to 2013 and 2014.

25% of the top 20 relievers and 36.7% of relievers ranked 21-50 according to RA9_WAR in 2010-2012 were top 50 relievers in 2013 and 2014. A number of top relievers according to RA9_WAR from 2010 to 2012 such as Sean Marshall, Jonny Venters, Jason Motte, Eric O’Flaherty, Jesse Crain, Rafael Betancourt, and Joel Hanrahan suffered injuries in 2013 or 2014 that significantly impacted their value. It is questionable whether the large amount of injuries for relievers ranked in the top 20 in this category is typical.

In addition, success rates weren’t impressive when using multiple measures to determine reliever quality. The five relievers ranked in the top 20 using each measure were Craig Kimbrel, Mariano Rivera, Sergio Romo, Mike Adams, and Darren Oliver. Kimbrel has remained excellent and is easily a top five reliever while Rivera retired after 2013, but Romo, Adams, and Oliver were all disappointments. Likewise, there were nine relievers ranked in the top 20 of two of the categories and in the top 50 in the other category. Four of those nine relievers ended up getting hurt and unsurprisingly only four of the relievers in that category were actually successful. Of the nine relievers who were ranked in the top 20 of two of the metrics without being ranked in the third, the only successful ones were Tyler Clippard and Joaquin Benoit.

The problem is that elite relievers have very little room for error. The top relievers had an ERA of about 2.00. If they throw 63 innings a season, then that means they can only allow 14 runs. If they end up allowing just eight more runs per season, then their ERA is closer to 3.14 and they are simply average. The difference between the best relievers and decent relievers is minuscule and can come down to a few good or bad breaks. At the same time those same extra seven runs are more important than the average run because elite relievers are usually used in the most crucial situations.

Some top relievers from 2010 to 2012 were still good in 2013 and 2014. But on the whole a top reliever from 2010-2012 was unlikely to be elite in 2013 and 2014. It doesn’t make sense to pay relievers for past performance because it isn’t likely that they will be able to repeat it. A team that has limited amounts of money should focus on either position players or starting pitching. Quality relief pitching is important, but there are so many variables involved that are outside the pitchers' control that even the best relievers can't consistently provide it.