Showing posts with label FIP. Show all posts
Showing posts with label FIP. Show all posts

15 June 2015

Miguel Gonzalez Is Still Cheating His FIP

In the off-season, a number of Camden Depot writers wrote articles discussing how Miguel Gonzalez has outperformed his FIP. This year has turned out to be no exception so far as he has an ERA of 3.33 and an FIP of 4.60 based primarily upon a .235 BABIP (2nd best out of 106) and an 82.2% LOB (7th best out of 106).

What is interesting is that this year he is outperforming based on performance against hitters when runs are in scoring position. He’s faced 54 batters in that situation and he’s allowed 8 walks (2 intentional), 4 home runs and caused 7 strikeouts. This is pretty awful because pitchers that give up more walks than strikeouts usually struggle and giving up a home run every ten batters on average is awful. He’s allowed seventeen fly balls and given up four home runs for an HR/FB of 23.5%. Simply put, those numbers are dreadful. Two of the remaining thirty-five batters have hit doubles which is pretty reasonable and shows that hitters aren’t struggling to hit for power against him when runners are in scoring position. The area where Miggy has excelled is that he’s only allowed one single with runners in scoring position. That’s absolutely incredible and has allowed him to put up a BABIP of only .083 with runners in scoring position. In addition, he has only given up one unearned run so it’s not like runners are scoring but not counting against him. Batters have been able to put up strong power numbers against him, but can’t seem to hit for average.

With runners in scoring position, Gonzo has a Hard% of 30.8% (75th out of 118 qualified pitchers), which indicates that twelve balls have been hit hard against him. His Med% of 51.3% (51st out of 118) indicates that twenty balls have been hit with medium force against him. His Soft% of 18% (68th out of 118) indicates that only seven balls have been hit softly against him. He isn’t causing weak contact against them, which might explain his low BABIP. Rather hitters are hitting the ball with adequate power but that hasn’t translated to them hitting singles.

ESPNs TruMedia tool claims that Gonzo has thrown 206 pitches with runners in scoring position of which batters swung at 106. Gonzo only had 19 called strikes out of 100 pitches where a batter didn’t swing. Of the 39 pitches put into play, only 8 (21%) were outside of the strike zone.  This seems to indicate that batters aren’t struggling against him because they’re swinging at pitches outside of the strike zone. It seems like they’re swinging at pitches that they’re able to hit.

Part of the reason why Gonzo is doing so well was touched upon by Ryan’s earlier article about Gonzo. Gonzo has thrown a total of 81 combined sliders and splitters with runners in scoring position, allowed 11 batters to put the balls in play and only allowed 1 hit total (a double) with those pitches. By comparison, Gonzo has thrown 13 curveballs with runners in scoring position and allowed both batters to put the ball into play to get hits, thrown 11 changeups and allowed 1 of 3 batters to put the ball into play to get hits and thrown 98 fastballs and given up 3 hits on 23 balls in play. Hitters are simply unable to convert when facing his slider or splitter and have struggled against his fastball. Ryan argued that Gonzo’s improved slider and splitter have allowed him to strike out more batters, but so far it also seems that they’ve helped him by causing a lot of unproductive contact.

The other part of the reason of why Gonzo is doing so well is that he has allowed a low but reasonable six line drives with batters in scoring position of which three were hits. He also has allowed 17 fly balls for which he’s given up four home runs and no other hits, 14 ground balls of which zero were hits and 2 bunts that also didn’t result in a hit. So far, batters have been able to hit fly balls out of the park with a staggering amount of success, but they’ve had limited success when hitting line drives and zero success when hitting fly balls that remain in the park or a ball hit on the ground. Either Miguel Gonzalez has found a way to have batters hit the ball towards fielders or has been the beneficiary of nearly flawless performance by his fielders with batters in scoring position.

It seems that it is unlikely that this will continue. Even if the slider and splitter have legitimately improved, it seems unlikely that batters will continue to struggle as much as they have against his fastball. If the defense has merely over performed, then it is unlikely that they’ll continue to do so to this extent. Then again, a number of writers have been saying for years that Gonzo can’t continue to be this successful and he’s done an excellent job of finding a way to outperform what his peripherals suggest his performance should look like. If some pitchers like Rick Porcello can find a way to mostly underachieve despite strong peripherals then it stands to reason that a few pitchers could find a way to overachieve despite weak peripherals. When it comes down to it, Gonzo has been good the past few seasons and is well on his way of having another strong season. Sometimes it’s just better to be lucky than good and no one is going to care about his FIP if he continues to perform at such a high level.

31 March 2015

FIP and the Ball-Strike Count II

In my previous post about FIP, I discussed how pitchers give up harder contact in a hitters count resulting in more extra base hits and weaker contact in a pitchers count resulting in more singles. This means that while BABIP doesn’t change much based on count, the effect of the hits based on the count does. I attempted to find metrics that might help predict which players could benefit from this but was unable to do so.

In this post, I decided to look at pitchers that gave up weaker than average contact and see how their ERA and FIPs compare to each other. From 2000-2013, I looked at all pitchers that gave up fewer than average doubles, triples and home runs while giving up a larger than average percentage of singles on balls in play. Presumably, pitchers that fit this profile give up weaker contact than pitchers that don’t and therefore should have a lower ERA than FIP. This test should help determine the impact of giving up weaker contact.

There were 117 pitchers that fit these criteria and many of them were ones that one might expect. For example, star closers such as Mariano Rivera, Craig Kimbrel, Joakim Soria, Ryan Cook, Jim Johnson (not including his terrible 2014), B.J Ryan and David Robertson were all on this list. So were guys like Rick Porcello, Brett Anderson, Brandon Webb, Chien-Ming Wang, Doug Fister and Derek Lowe. These pitchers are well known for giving up a lot of ground balls and allowing only weak contact. A list of the entire 117 pitchers can be found here.

It should come as no surprise that the pitchers on this list record more saves than their average usage would suggest. These pitchers threw only 9% of total innings while recording 20.8% of total saves. It makes sense that pitchers that can avoid giving up hard contact end up being used as closers.
However, there are some surprising results when we compare their ERAs to their FIPs. Only 58 of the 117 pitchers on this list actually have an ERA lower than their FIP. The mean ERA is 3.73 while the mean FIP is 3.78 or about a difference of .05. This is minimal and unimportant. This suggests that even pitchers that end up allowing weaker than average contact still have an ERA that’s similar to their FIP.  It appears that despite the fact that FIP doesn’t differentiate between a single, double or triple the stat still accurately describes performance.

The pitchers that primarily are able to outperform their FIP are those that are able to avoid giving up any type of hits whether they’re singles or for extra bases. There have been 26 pitchers that give up fewer 1B, 2B, 3B and HR than the 35th percentile. They have an average ERA of 3.03 and an average FIP of 3.55. But then again, it’s questionable whether that should be attributed to pitcher skill or to defense.

This doesn’t mean that FIP is necessarily 100% accurate. Some have proposed that Chris Tillman has extra value not measured via traditional statistics because he is able to keep opposing batters close to the bag and therefore prevents steals, prevents runners from advancing multiple bases on a hit and creates more double plays. But it is interesting to see that FIP “works” even in a situation where we’d expect it to fail.

This analysis indicates that pitchers do have some effect on how hard their pitches are hit and therefore whether impact the likelihood of a batter getting an extra base hit. Some pitchers appear to be better than average at preventing hard contact than other pitchers while most pitchers give up weaker contact in favorable pitch counts. I have been unable to determine an impartial way of determining which pitchers should be expected to give up weaker contact than their peers but it appears that someone using the eye test or knowledge about pitchers can predict this with reasonable accuracy. However, FIP appears to be accurate even when looking at pitchers that allow weaker contact than their peers. I am not quite sure how this can be the case but facts are facts. This possibly indicates that FIP is a better metric than its detractors suspect and suggests that people shouldn’t necessarily dismiss it simply because they find things that it doesn’t measure.

17 March 2015

FIP and the Ball-Strike Count

Voros McCracken argued in 2001 that pitchers have little ability to prevent balls in the field of play from becoming hits and therefore a pitcher’s ability to record strikeouts, avoid walks, and limit home runs separates good pitchers from bad ones. He argues that one reason for this is that there’s a significant lack of year-to-year correlation in BABIP.

One of the questions I’ve had about this is what this means for ball-strike counts. If pitchers have little impact on balls in play then presumably they should give up a similar percentage of singles, doubles, and triples on a 0-2 count than on a 3-0. This seems unlikely and unintuitive because a pitcher can’t be as selective in a 3-0 count than in a 0-2 count. Using data from Retrosheet, I built a data file with all at-bats from 2000-2013 that ended up in either a single, double, triple, home run, or out in play and determined the count when the ball was hit into play.  Below is the data.


These results do partly support FIP because as the count becomes more hitter friendly it becomes considerably more likely that the ball will be hit out of the park. Hitters hit a home run four times more often on a 3-0 count than on a 0-2 count. The idea behind FIP is that pitchers have more control over home runs than over balls in play and these results certainly back that idea up.

However, the results also show that triples and doubles are more likely in a hitter's count and singles are more likely in a pitcher's count. I think this chart is showing that pitchers are giving up harder hit balls on hitter's counts than in pitcher's counts. While I wouldn’t have expected these results, they do make sense because it’s easier to hit a lucky single than a double, triple, or home run.

The BABIP stat is very interesting because while it indicates that hitters do have a better BABIP in a hitter's count than in a pitcher's count, the impact is minimal. If one was just looking at BABIP and didn’t look at the type of hits then the differences would appear to be minor. This means that a lack of significant year-to-year BABIP is mostly irrelevant. The reason why pitchers have better results on balls in play in a pitcher's count rather than a hitter's count is because they’re giving up fewer extra-base hits and not fewer hits overall. Presumably there should be some value in giving up singles rather than doubles or triples.

It appears that this data does indicate that pitchers do better in pitcher's counts than in hitter's counts. The next question is whether certain types of pitchers are more likely to allow contact in pitchers counts rather than hitter's counts or whether they’re able to give up lighter contact. Basically, if we can identify pitcher groups that give up weaker contact than other groups then this would potentially show a flaw with FIP. Alternatively, if all pitchers allow contact at the same rate for each ball-strike count then what I found above may be interesting but is purely academic.

I’m not sure which metrics are the best to predict ball-strike count or the type of contact. The two metrics that I did use to see whether this is the case was K% and LD%. If pitchers with a high K% strike out a lot of batters then it seems possible that they’ll face more pitcher-friendly counts than those that have low K%. If so, logic would indicate that they’d give up more contact in pitcher-friendly counts. Likewise, pitchers that give up fewer line drives than average should probably give up lighter contact than those that give up more line drives than average. I’ll start by looking at K% and the data is below.


The data is interesting. It indicates that pitchers with above-average K% do considerably better in hitter's counts than those with below-average K%, and it would be interesting to figure out why that’s the case. It’s possible that it's primarily due to small sample size. However, all in all, pitchers with below average K% give up .03% more singles, .11% more doubles, .01% more triples, and .04% fewer home runs. Furthermore, this next chart shows how often they give up contact for a given ball-strike count.


Pitchers with below average strikeout rates do give up slightly more contact in hitter-friendly counts than in pitcher-friendly counts. The problem is that the difference is minimal. Simply put, this metric doesn’t show that pitchers with high K% either give up lighter contact or pitch in significantly more pitcher-friendly counts than those with low K%. Next up is LD% rates and below is the data.


Basically, the data show that pitchers with above-average LD% rates (lower than average) give up fewer singles, doubles, and triples than those with below-average LD% rates while allowing the same amount of home runs. This suggests that FIP may be improved by looking at the type of contact that a pitcher gives up but ultimately isn’t helpful for my purposes. The chart below shows that pitchers with above average and below average line drive rates allow contact at roughly similar rates for each unique ball-strike count.


Since I already looked at BABIP before realizing that it wouldn’t be helpful, I suppose it makes sense to discuss that quickly as well. Below is a chart with data.


Basically, it also shows that pitchers with good BABIPs allow roughly 18% fewer singles and doubles as well as 12% fewer triples than those with bad BABIPs. Pitchers with good and bad BABIPs also allow contact at roughly similar rates for each unique ball-strike count.

At this point, I’ve shown that pitchers have control over primarily their home run rates but also balls in play based on the count. However, I haven’t been able to show that there are certain types of pitchers that are able to be successful because they are able to minimize contact allowed in pitcher's counts. This could be for any of at least three reasons.

The first reason is that I simply could have picked poor metrics. Just because the metrics that I used to see whether certain pitchers minimize contact in pitcher's counts didn’t do that doesn’t mean that other metrics won’t be more successful.

The second reason is that this could be a fringe skill. It may be the case that only a few pitchers are able to consistently avoid hitter's counts and therefore give up fewer extra-base hits. If so, perhaps looking at only 10% of the population instead of 50% of the population will provide better results.

The third possibility is that pitchers aren’t able to control when they allow contact and therefore noting that avoiding hitter's counts reduces extra-base hits is merely academic.

Next week, I’ll look at this question with a different metric and see whether that changes the results.

30 January 2015

Should Teams Focus on Run Prevention?

In my previous article, I argued that the Orioles should avoid spending money on any pitching free agents in order to spend resources on position players that can provide offense and defense. My argument was that excellent fielders can turn good pitchers into excellent pitchers and so on and so forth.  But it’s also possible that excellent fielders combined with excellent pitchers can prevent enough runs that even a mediocre offense can score enough to win. Does it make sense for teams to primarily focus on run prevention?

In order to answer this question I looked at all starting pitchers that threw at least 100 innings for a given team from 1935-2014 and put them into quartiles based on FIP. Then I looked at all teams from 1935-2014 and put them into quartiles based on their Fangraphs fielding metric. 

The first chart shows fielder and pitcher performance order by fielding ratings. E/F stands for ERA/FIP while RA9/F stands for RA9/FIP. I include these stats to measure the impact of defense based on a proportion rather than just a total. Here it is below.




This chart shows that the value of having a good defense decreases when you have good starters. For good fielding teams the difference between ERA and FIP is greatest for the worst pitchers. For poor fielding teams this is still the case but the difference is minimal. It is still better to have excellent pitching than having average pitching but teams have limited resources and may not be able to easily afford top pitching.

The second chart shows pitcher and fielder performance ordered by pitcher ratings.



This chart shows that good pitchers (those in the second quartile) that have the best fielders prevent more runs than the best pitchers with the worst fielders. Bad pitchers that have the best fielders prevent as many runs as good pitchers with bad fielders. Bad pitchers with the best fielders behind them give up only 3 more runs per 200 innings than good pitchers with good fielding. The best pitchers do prevent more runs on average than worse pitchers but good fielding is able to significantly limit the gap.

The chart also shows that even the best fielders are unable to do much for the worst pitchers. It is true that the worst pitchers with the best fielders have an RA9 nearly .7 points lower than the worst pitchers with the worst fielders. It’s also the case that teams with the worst pitchers and best fielders give up 4.7 runs on average per 9 innings and are still ineffective.

This next chart shows how it worked in practice for the 2014 Orioles.



The chart shows that the Orioles have seen this happen first hand. The only starter we had that had an above average FIP was Gausman and he was the only starter who underperformed his FIP. The other starters had an FIP that was either in the third of fourth quartile and mostly outperformed what their FIP suggested. One could argue that the Orioles actually did have a rotation of below-average starters and were able to succeed due to superior fielding. It’s not that Chris Tillman or Wei-Yin Chen are good necessarily but that they have an excellent defense supporting them.

Miguel Gonzalez outperformed his FIP by a substantial amount. There have been 8,487 starters that have thrown 100 innings for one team in a given year from 1935 to 2014. In 2014, Miguel Gonzalez has the tenth highest E-F and the fifth highest RA9-F. Players with similar seasons include Jorge Sosa in 2005 and J.A Happ in 2009. It seems likely that Miguel Gonzalez should expect significant regression if used as a starter next year.

It appears reasonable to argue that the Orioles have been able to prevent runs because they have a large number of decent starters that are able to succeed because they have a strong defense behind them. They may not have the best pitchers in MLB but good defense can fix a lot of problems.

In order to continue being successful, the Orioles need to consider pulling a trick out of the Rays playbook. The Rays typically offer pitching prospects long term deals before they are proven in order to save money and then trade them a few years before they become free agents. Of the Rays 10 starters that have thrown the most innings for the Rays from 2005-2014, they traded seven, had two get injured and currently control only Alex Cobb.  In contrast, the Rays only traded four of their ten offensive players with the most PAs from 2005 to 2014. Three of those four were traded this offseason in order to save money.

The Rays were able to do this because their fielding has been ranked #1 by Fangraphs from 2008 to 2014. Having good fielders supporting their pitchers allows them to risk giving guaranteed money to unproven pitchers and risk trading their proven pitchers to other teams.  They weren’t that aggressive when it came to position players because they needed to focus on offense and fielding in order to win.

The Orioles should consider a similar strategy.  Signing Gausman and especially Bundy to long-term contracts later this year will save the club money if they are successful. Trading a starter like Wei-Yin Chen or Chris Tillman could potentially net a reasonable return while solving our rotation questions.

Excellent fielding is more helpful for decent pitching than excellent pitching. That means that teams shouldn’t focus on trying to find the best fielders and starters at the cost of offense. Spending money in free agency to improve pitching is a luxury that a team with limited resources simply can’t afford. The best plan is to use most available resources for offense and fielding.

26 January 2015

Chris Tillman: Good Pitcher, Not an Ace

Since Chris Tillman became a permanent member of the Baltimore Orioles rotation in July 2012, one could argue that he’s been the team’s best starting pitcher, albeit on a pitching staff that has been less than stellar during that time.   Since the 2012 season, Tillman has pitched about 500 innings, with an ERA of 3.42, which is better than any other Orioles starter by almost a half of a run.  A difference of a half run of ERA isn’t trivial over the course of the entire season.  Assuming 200 innings pitched, it adds up to 11 or 12 additional runs prevented.
Chris Tillman (photo via Keith Allison)

The topic of whether Chris Tillman is an ace isn’t a new one.  It’s been covered elsewhere, and has been covered here at Camden Depot as well.  Our own Jon Shepherd took on the topic in 2013 after receiving a reader email and then again towards the end of the 2014 season in a post he wrote for MASN.  In his previous posts, Jon first attempts to define an “ace”.  While he used slightly different methods, in both instances he arrived at the same place: in any given year the league has roughly 10 “aces”.  He then follows that discussion explaining why he believes that Chris Tillman does not qualify as an ace based on his previously defined designation.

While I in no way disagree with the conclusion of the two aforementioned articles, both of them generally used Fangraph’s version of WAR (and the components that contribute to it) as the basis for measuring whether Tillman should be considered an ace.  However, fWAR, FIP, and the statistics that reward them have never been Tillman’s strong suit.  He isn’t especially great at striking batters out (of qualified starting pitchers, he ranks 47th in K% since 2012), limiting walks (ranked 82nd in BB%), and doesn’t do a great job of keeping the ball in the yard (ranked 105th in HR/9).  It’s no wonder that despite that 3.42 ERA mentioned earlier, Tillman has only been worth 5.6 fWAR since 2012, more than a win less than Wei-Yin Chen in only slightly less innings.

However, what Tillman has done well since 2012 is prevent runs, as noted by his 3.42 ERA during that time.  Run prevention plays a much bigger role in calculating Baseball-Reference’s version of WAR, so it’s not a surprise he’s been worth almost 3 more wins by their standards compared to Fangraphs over the same period of time.  It’s believed that certain pitchers have a skill that FIP is unable to capture, and will therefore post ERA’s that exceed their FIP on a regular basis (Matt Cain of the San Francisco Giants is the most recent example).  Since 2012, Tillman’s ERA has been better than his FIP by AT LEAST 0.67 runs each year (with a maximum difference of 1.32 runs during 2012).  Could Tillman be one of those rare pitchers that FIP just doesn’t understand?  It’s certainly possible, but it’s also probably too soon to tell (one could say the same for Miguel Gonzalez).

Hypothetically, let’s assume that Tillman does have the ability to regularly outperform his FIP.  If so, can he be considered an ace based solely on his ability to prevent runs?  In order to determine this, I looked at ERA+, which takes into account ballpark effects and the offensive environment of the era.  As Jon mentions in his articles, people can have many different views as to what constitutes an ace.  Because of that, I looked at the average, maximum, and minimum ERA+ values for the 10th, 20th, and 30th best pitchers from 1961 to 2014.  


I then compared those values to Tillman’s ERA+ numbers since 2012 to see where Tillman fits in, if at all.

*From 2012, when Tillman only pitched 86 innings
It’s only two plus years of data for Tillman, but if you squint and use the most broad definition of what is considered to be an ace (one of the top 30 pitchers in baseball), one could make an argument that Chris Tillman is an ace when considering run prevention only (he’s obviously doesn’t fit in the more stringent ace categories).  However, I personally don’t consider a top 30 pitcher to be an ace (I’m in the “10-15 aces” camp), and even though Tillman could be viewed as a top 30 pitcher in baseball in terms of run prevention, the fact that he isn’t close to being a top 30 pitcher in terms of FIP and fWAR (as detailed in the previously linked to posts) further diminishes what little case Chris Tillman already has as an “ace”.

Just because Chris Tillman isn’t technically a traditional ace, does not mean he isn’t a good pitcher.  If you prefer FIP, he’s likely a number 3 or 4 pitcher on a good team.  If you’re more partial to strictly run prevention, he could probably be considered more of a number 2 or 3 pitcher.  Either way, he should provide decent value to the Orioles, even as he enters his arbitration years (Tillman and the Orioles avoided his first arbitration hearing last week by agreeing to a $4.315 million contract).  Of course, how much value Tillman provides will depend on whether he is in fact one of those rare pitchers who can consistently outperform his FIP, and remain on the fringes of being a “run prevention ace”.

22 January 2015

How the Orioles Broke FIP

One could argue that the golden age of the Orioles Franchise was from 1960-1985. Over those 26 years, the Orioles won three World Series, lost another three World Series and were knocked out in the ALCS twice. The Orioles had twenty-four winning seasons over that time frame including eighteen consecutive winning seasons from 1968-1985. They had a cumulative record of 2374-1749 or a winning percentage of 57.6% which was easily the best winning percentage in baseball over that period. The next closest was the Yankees with a winning percentage of 55.6%. Outside of that time frame, the Orioles haven’t made it back to the World Series (although the St Louis Browns did in 1944) and have made it to the playoffs just four times. All in all, it was a pretty good run.

When I was looking at some of the numbers from that 26 year time frame, I noted a few interesting things. Our offense was one of the best in the majors but averaged the third most runs per game. The Red Sox led the majors during that time frame and scored 4.53 R/G compared to our 4.35 R/G. Then again, the Orioles won nearly 240 more games than them over that period or more than 9 wins per season.

The Orioles had an FIP of 3.65 over that 26 year period which tied for ninth in the majors. This is decent but indicates that their pitching was not as good as one would expect from a team that was dominant for 26 years. If the Orioles offense was great but not elite while the pitching was merely good then how were the Orioles so successful over that period of time?

Over that 26 year period the Orioles outperformed their FIP by .29 points. The next best team was the Yankees who outperformed their FIP by .17 points.  In addition, their RA_9 was just .08 points larger than their FIP. The Yankees were the next best club but their RA_9 was .27 points larger than their FIP or nearly three times the difference. In absolute terms, the Orioles allowed only 348 runs more than their FIP suggested.  The next closest team was the Blue Jays who allowed 581 runs more than their FIP suggested. Then again, the Blue Jays gave up that many runs in 12,495 innings while the Orioles gave up that many runs in 37,163. As a result, despite having the ninth lowest FIP, the Orioles allowed the second fewest runs per game in the majors over that 26 year period. The data suggest that when the Orioles were elite it was because they had a strong offense and were giving up fewer runs that their FIP suggested. In essence, the Orioles broke FIP.

Last week I noted that even an elite defense should only be expected to outperform their FIP by about .20 points. So how were the Orioles able to outperform their FIP by nearly .3 points? According to Fangraphs fielding metric, the Orioles defense was valued at 1276 runs from 1960 to 1985. The next best team was the New York Yankees and their defense was valued at only 498 runs over that time period. The #2 through #4 teams defensively were worth 1239 runs from 1960 to 1985. The Orioles defense was about two and a half times better than the second best defense in the time frame or by about thirty runs per year. The reason why the Orioles outperformed their FIP by the extent that they did is because their defense wasn't merely elite. It was legendary.

A closer look at the defensive results suggests that the Orioles weren’t elite defensively at each position. Rather, they focused on having strong defensive players at second base (#2), shortstop (#1), third base (#1) and center field (#1) but weren’t particularly good defensively at first base, catcher, right field or left field. All of this makes perfect sense because it is common baseball knowledge that a team needs good defensive players at those positions.  Meanwhile, the Orioles weren’t particularly good offensively at center field (#11), DH (#7), third base (#10) and catcher (#17). They relied on offense from first base, second base, shortstop (it’s more that their players at this position weren’t as bad as most teams) and outfield.

If the Orioles were able to be elite from 1960 to 1985 at least in part due to their excellent defense than it makes sense for the Orioles to try and focus on what once made them great. Excellent defense makes good pitching look like its elite and therefore the Orioles shouldn’t be focusing on trying to find the best pitchers. Rather, the Orioles need to focus on doing the following.

They need to avoid spending big dollars on any pitching free agents. Jim Palmer won three Cy Young awards, had eight twenty win seasons and made it into the Hall of Fame on his first year eligible with 92.6% of the vote. He also benefited greatly due to having an elite defense behind him. Is he a Hall of Famer if he was on the Cubs instead of the Orioles? Dave McNally and Mike Cuellar also had strong ERAs and only decent FIPs. Were they great pitchers or simply lucky to play behind a dominant defense? It seems reasonable to conclude that an elite defense turns great pitchers into Hall of Fame caliber pitchers, good pitchers into great pitchers and decent pitchers into good pitchers. If you don’t have enough resources to focus on everything than ignore pitching and trust your defense to make your pitchers look good.

That allows most of our resources to be devoted on building a strong offense and an elite defense. They need at least two elite players that are strong both offensively and defensively at either second base, shortstop, third base and center field. They need to find another two defensive wizards that can fill the other two positions. And they need three players that are strong offensively that can play at either first base, left field, right field and DH.

What’s interesting is that the Orioles are loosely following this plan. The Orioles have potential elite offensive and defensive talent at third base and shortstop. Schoop is excellent defensively at second base while Adam Jones has elite offensive ability at center field. The Orioles have potential above average offensive talents at first base, catcher and right field. The Orioles would strongly benefit from having an excellent center field option that can both hit and field that could push Adam Jones to right field. It would also be nice to add another platoon bat that can team up with Delmon Young to hit right handed pitching. Still, the Orioles are close to having the offensive talent to be successful.

The Orioles have wasted some resources on pitchers like Ubaldo Jimenez. But for the most part their rotation is filled with pitchers that may not seem to be special but have a knack for over performing what is expected from them. They may not be as successful on a team with poor fielders but they're on the team that they're on. More importantly, most of their pitchers are cheap which makes it possible to spend money on the more important position players.

What if the Orioles decided that what they’ve been doing wasn’t working and therefore decided to go back to what has been successful in the past?  This team has a lot of similarities to the Orioles teams in 1960-1985 and with two playoff appearances in the last three years the plan seems to be working.

14 January 2015

Orioles' Defense Isn't as Good as You Might Think

On Monday, Ryan wrote an article discussing how the Orioles' pitching isn't as good as you think. Commenters reasonably argued that FIP consistently penalizes teams with good defense and therefore using that metric will underestimate the Orioles' pitching. However, this does beg a few questions. Do teams with good defenses usually have an ERA lower than what their FIP suggests? If so, how large is the average impact? Steamer projects the Orioles' pitching staff to have an ERA of 4.04 and a FIP of 4.30, or that the O's defense will prevent 42 runs over an entire season. Does this underestimate or overestimate the projected impact of the defense?

In order to answer this question, I looked at each team from 2000-2014 and determined their ERA, FIP, and fielding score defined by Fangraphs. I then split them into quintiles based on fielding score. I didn’t use Fangraphs' defense metric because that punishes AL teams for having a DH. For this kind of analysis it simply isn’t as accurate as the fielding score metric. I also included the 2014 Orioles as their own special category to see whether they were an outlier. This chart shows the results:


Group ERA FIP E-F Difference Between ERA and FIP Fielding
1_Worst 4.55 4.32 0.24 38.32 -48.64
2_Bad 4.36 4.28 0.08 12.74 -17.05
3_Average 4.29 4.25 0.04 7.20 -0.27
4_Good 4.03 4.17 -0.14 -22.84 18.92
5_Best 4.08 4.29 -0.21 -34.69 47.93
2014 Os 3.44 3.96 -0.52 -84.24 56.40

There does appear to be a relationship between teams’ fielding scores and whether or not they do better than their FIP suggests. The teams with the worst defense had an ERA about .24 points larger than their FIP, which meant they allowed over 38 runs per year more than their FIP suggested. Teams with the best defenses had a FIP that was nearly .21 more than their ERA, which meant they allowed nearly 35 fewer runs than their FIP suggests. Fangraphs' fielding metric successfully predicts which teams will do better or worse than their FIP.

However, one would expect the difference between ERA and FIP to be similar to the average fielding score for each group.  As the number of teams sampled increases, it should be expected that potential issues such as sequencing and luck are less of a relevant factor. However, as the fielding score increases (whether negative or positive) the difference between it and a teams’ ERA-FIP becomes more pronounced. For example, the worst clubs have an average fielding score of -49 runs but the difference between their ERA and FIP is only .24 runs or about -38 runs. The best clubs have an average fielding score of 48 runs but a difference between ERA and FIP of only .21 runs or about 35 runs. This suggests that either the Fangraphs' fielding metric inflates the value of defense or that defense has diminishing returns.

The 2014 Orioles had a stunning .52 run difference between FIP and ERA. This is more than twice as high as the difference between FIP and ERA for teams that have even the best fielding scores. In fact, the 2014 Orioles had the third-largest difference between ERA and FIP from 2005-2014. This suggests that the Orioles' defense could be elite in 2015 and still wouldn’t be expected to outperform their FIP by such a drastic amount. The Orioles' defense does explain why they are better than their FIP but not why they were better by nearly 85 runs. It seems that the difference should be closer to 40 runs.

The other test that I did was model the difference using a regression between ERA and FIP for each team from 2000 to 2014 based on Fangraphs' fielding metric.  My results were statistically significant with an R^2 of .3836 or an R of .6194. When I tested all data from 1935 to 2014, my results were statistically significant with an R^2 of .4365 (R of .66).  These are moderate to high correlations and suggest that Fangraphs' fielding metric can be used to predict which pitching staffs will do better than their FIPs suggest.  However, it also suggests that there are other relevant variables and that just using Fangraphs' fielding metric may not consider all relevant factors.  Furthermore, it is highly unlikely that a sample consisting of nearly 2,000 seasons would see an impact from uncontrollable factors such as hit sequencing. As Beyond the Box Score notes, it is likely that disparities between ERA and FIP could be impacted by pitching performance as well as fielding. This could potentially further explain why the 2014 Orioles had such a large difference between their ERA and FIP and could potentially suggest even more inflation in Fangraphs' fielding metric.

Having an excellent defense means that pitchers will give up potentially 30 to 40 runs fewer than their FIP suggests, which is roughly the impact that Steamer projects. This doesn’t explain why the 2014 Orioles were able to allow 85 fewer runs than their FIP suggests, and suggests that unless other factors can explain this discrepancy we should expect significant regression in 2015.

12 April 2012

How Good is an NL Ace?: Mean Performance of Pitchers by Slot

There is a series of articles by Jack Sackman that you can find here.  It is an idea I found interesting an often use when I describe pitchers as a certain type of slot pitcher.  I think in common use a person referring to a guy as a one slot pitcher is more or less actually saying that the guy is a one slot pitcher on a first division team.  In other words, an ace on one of the ten best teams in baseball.  In this series of posts, I plan on going through each division and describing what each slot means and how that relates to teams.
AL East | Central | West
NL East | Central | West
NL Summary of Slots

In this post we will go through and look at four team FIP performances for each slot: median, first division cut off, best, and worst.  The following relates to numbers produced in 2011.

Slot 1
An NL pitcher at this slot could be described as:

Median 3.15
67th 3.24
33rd 3.05
The Nationals' Jordan Zimmerman (3.16 FIP) is your typical ace pitcher.  Philadelphia Phillie Cole Hamels (3.05 FIP) would be the threshold first division ace and Cardinal Jaime Garcia (3.23 FIP) would be the closest to a bottom rung ace.  In 2011, the Phillies had the best ace (Roy Halladay 32g 2.20 FIP) and the Astros were the worst (Lucas Harrell 2g 3.27 FIP; Bud Norris 30g 4.02 FIP).

Slot 2
An NL pitcher at this slot could be described as:

Median 3.62
67th 3.79
33rd 3.48
The average Slot 2 pitcher would be the Brewers Yovani Gallardo (3.59 FIP).  Reds' Ace Johnny Cueto (3.45 FIP) is the threshold first division second slot.  New Yankee Hiroki Kuroda (3.78 FIP) would qualify as a bottom threshold second slot pitcher.  The team with the best slot 2 performance are the Philles again (Cliff Lee 32g 2.60 FIP) and the worst was the Pirates (Paul Maholm 23g 3.78; Brad Lincoln 8g 3.88 FIP; Aaron Thompson 1g 3.95 FIP).

Slot 3
An NL pitcher at this slot could be described as:

Median 3.91
67th 4.05
33rd 3.73
Your typical three is the Cubs' Ryan Dempster (3.91 FIP).  The first division gate keeper is the Brewers Shawn Marcum (3.73 FIP).  Jason Marquis (4.05 FIP) would be the bottom rung three man.  The Phillies again have the best performance for the slot (Cole Hamels 31g 3.00 FIP; Vance Worley 1g 3.24 FIP) and the worst was, once again, the Pirates (Jeff Karstens 26g 4.29 FIP; James McDonald 6g 4.68 FIP).

Slot 4
An NL pitcher here can be described as:

Median 4.23
67th 4.55
33rd 4.08
Jhoulys Chacin (4.23 FIP) of the Rockies is the pitcher that embodies the meaning of the fourth slot.  The Brewers' Chris Narveson (4.06 FIP) would be your threshold 4 man and the Phillies' Kyle Kendrick (4.55 FIP) would be your lower tier line.  The Phillies once again set the tone here with the best 4 slot squad (Vance Worley 20g 3.24 FIP; Roy Oswalt 12g 3.44 FIP) and the worst was the Reds (Mike Leake 7g 4.21 FIP; Sam LeCure 4g 4.57 FIP; Edison Volquez 20g 5.29 FIP; Bronson Arroyo 1g 5.71 FIP).

Slot 5
An NL pitcher here can be described as:

Median 4.64
67th 5.27
33rd 4.41
The median 5 slot pitcher would be the Mets' Dillon Gee (4.65 FIP).  Your first division fiver was the fellow Met Mike Pelfrey (4.47 FIP) and the bottom third gate keeper was the Reds' Edison Volquez (5.29 FIP).  The Phillies sweep the slots (Roy Oswalt 11g 3.44FIP; Joe Blanton 8g 3.55 FIP; Kyle Kendrick 13g 4.75 FIP) in better fashion than the Orioles who rated last across the board. The team with the worst back end performance in the NL was the Diamondbacks (Joe Saunders 5g 4.78; Wade Miley 7g 4.79 FIP; Micah Owings 4g 4.85 FIP; Jason Marquis 3g 6.91; Barry Enright 7g 6.98 FIP; Armando Galarraga 6g 7.29 FIP).

AL Average Rotation
1 - Jordan Zimmerman, Nationals
2 - Yovani Gallardo, Brewers
3 - Ryan Dempster, Cubs
4 - Jhoulys Chacin, Rockies
5 - Dillon Gee, Mets

AL First Division Threshold Rotation
1 - Cole Hamels, Phillies
2 - Johnny Cueto, Reds
3 - Shawn Marcum, Brewers
4 - Chris Narveson, Brewers
5 - Mike Pelfrey, Mets

02 April 2012

AL Team FIP by Pitching Postion

Keep an eye on the y-axis.  It changes from one graph to the next.  Also, the order of the teams change as well except for the Orioles who ranked as having the worst pitching by slot for every slot.

First Slot


Second Slot


Third Slot


Fourth Slot


Fifth Slot


As a little extra...here is the Orioles xWAR vs the World

Click to Enlarge