Report Reflection, Part 2: Types of Evidence for Sociology and Their Limitations
I first drafted this post a little over three weeks ago. My original plan was to release this post and my other post on the Friday when the report went live, but a decision was made to let the report do the talking, and so I held off to avoid distracting attention from a very carefully written report. Since initially I was going to release my internal report two weeks after the first report, I thought it would be fine. But now my report won’t be out until July (or maybe August), so it’s no longer fine, especially since the report is spurring all sorts of discussions on things I’ve been working on for about ten months and I don’t want to be left out!
However, there’s a danger in releasing this post, which is that some folks will think it is the extent of the internal report. So, for the record, my internal report on sociology, summarizing all my research, was 100 pages. It later (after completion in February) grew to 120 pages as more work came out or I learned about other pieces; I’m now working on cutting it back down after circulating it for feedback/informal peer review. Also, even more stuff has come out since I drafted this post and I’ve not had time to add it in. So, this post is not a summary, just some highlights of the types of things I discuss in my report, and what’s not in my report, that I’m most interested in. And I absolutely hate “tell, not show” writing, but that’s basically what I’m going to do here; I cite to some studies, but I’m not actually even really giving my findings, just the headlines. So stay tuned for the full actual report.
My task on the Commission was to evaluate scholarly standards in sociology. I produced an internal report that did just that, which then informed our general Report. My report primarily relied on others’ published research and commentary on these issues. (Sociology was by far the most studied field, so my job was mostly to search and synthesize.) There is other quantitative data that I wasn’t in charge of that will come out separately, later, and that I’m really looking forward to being public because it’s fascinating and systematic and large scale, and it asks questions I wouldn’t have thought to ask. Since submitting my report, I’ve also started other research projects to try to get at some of the questions I couldn’t answer based on the extant research.
In this post, I give a little taste of my report’s findings, but I focus more on the tensions, challenges, and limitations with the available evidence. This is not a summary, but more a methodological reflection.
Heterogeneity in Metrics
As I mentioned previously (and as the Report notes and as I try to say every time I discuss our findings), fields vary in terms of the magnitude of the problems they face. Within each field as well, there’s heterogeneity. But the problem also manifests to different degrees in different metrics. There is clearly good research produced across all the fields examined. There’s also a clear commitment to politics and an explicit rejection of certain scholarly standards many folks take as bedrock. Both can be true. But this actually presents an interesting methodological problem: how do you generalize about the scholarly standards of a whole field when it’s heterogeneous? My solution, which I’ve discussed before, is to be clear about the different metrics and the different findings associated with them. In light of all the evidence, what I can say is: scholarly standards are clearly corrupted, but we can’t say how much this is manifesting in worse research output, just skewed and biased* output.
*How skewed, how biased? Well, that depends and varies. In my view, though, the number itself doesn’t matter—bias in 10% of work would be too much, and we’re very likely past that point. What matters is you don’t always know which 10% because traditional markers—a job in a top program, a journal pub in a top journal, an award-winning article or book—are no longer indicators of quality because sometimes obviously low-quality work gets through. And then there’s the dark figure of how much good work is getting held back because it goes against the political grain.
To emphasize again, a key finding was the strength of the evidence varies by metric sort of inversely to the importance of the metric.
In sociology (as in other fields), the most superficial (but visible) metrics (explicitly activist or social justice professional conference themes and presidential addresses, public political statements from departments and professional associations) tend to show the strongest evidence of politicization supplanting a knowledge-seeking outlook. But these superficial indicators may not be representative of sociologists’ views (as a group) or actually tell us about the state of research produced. Indeed, it is an empirical question to what extent these high-profile episodes shape scholars’ views of the field. (Given some changes in the literature around the same time, it’s plausible they had an effect, but it’s also confounded because many things changed around the same time as some of the earliest, most explicit, and most discussed statements. Additionally, anecdotally, I certainly saw folks use some presidential addresses to justify moves that I don’t think they would have gotten away with before. So I think it had an effect, we just don’t know what size an effect or whether it was the dominant factor amidst multiple factors.)
The most substantive metrics (those relating to the production and reception of research) are harder to find because they are labor-intensive (e.g., careful reviews of the quality of articles published in top journals) or the requisite data isn’t readily available (confidential T&P or hiring proceedings). We do have LLM studies, but they only measure certain things, not the full range of quality checks.
Most of the research is in between: there is decent data about what is actually produced topically and about scholarly norms, and both are actually pretty strong on the politicization question. For me, these data are key, but they are adjacent to what I most want to know: the quality of specific output, especially the relationship between politicization and quality.
Lingering Questions
Let me jump to what I don’t know (or at least questions I had when I was done with my report back in February and that I’m now working on [or will resume when I’m done with revising my report]):
While there are clear negative consequences to the political biases in the field, we do not know to what extent this manifests in the quality of research produced. Where we do have some data on quality, it’s fieldwide and not stratified by journal prestige. But 90% of everything is shit, so what does this tell us. (The overtime analyses are actually helpful on that front, but even so, the objection remains.) We can point to individual articles that probably shouldn’t have been published in those top venues (and likely wouldn’t have a decade earlier). My subsequent (and still preliminary) research (not in the report) is showing that even in top journals, we’re seeing on average declines in quality, including previously common expectations about what should be in an article. But none of this is sufficiently systematic to be able to say how widespread these issues are. More research is needed. We also have good quantitative findings about topics covered, which show pretty widespread patterns in blindspots (even in top-tier journals), but again we don’t know how that relates to quality of the remaining topics. Some folks argue it has quality impacts, but we just don’t have good data to know for sure.
Is the commitment to activism, especially in combination with the rejection of scientific values like neutrality and (aspiring to) objectivity, consistent across the field or more common in/mostly restricted to lower-tier departments? (My top-tier colleagues claim it is. I think they are wrong. But we just don’t know.) Survey data on opposition to objectivity and a commitment to politics in research is pretty widespread in its coverage, but we don’t have survey answers stratified by program status (even though the surveys are usually R1 faculty, that’s still a huge mix). There are some obvious cases of individual faculty in top-tier programs doing questionable work, but identifying high-profile individuals (which I don’t do in the report) is not the same as systematic representation in good programs. In theory, good programs shouldn’t hire questionable scholars at all, so any examples are pretty bad, but we don’t know empirically what sort of damage they cause or how widespread they are.
These are my two most burning questions. I have others, but they are secondary and mostly related to additional areas where more research is needed (we have a lot of high-profile examples of horrifying situations, but only the worst cases get publicized and written about, so again questions of representativeness and about chilling effects remain open, which I discuss below).
A Survey of the Evidence
In recent years (and at times even going back several decades), the American Sociological Association, the field’s primary professional association, has been explicit in its commitment to social justice advocacy and activism in its conference themes (going back more than a decade), presidential addresses, and public statements on political topics such as the war in Gaza. This is the superficial, high-profile evidence that tells us about the state of the field, but not the state of regular research. However, I would argue it’s also important because it legitimizes what sorts of research is appropriate and encouraged by powerful actors in the field—literally our professional association—who are also telling us what it means to be a sociologist. As a proponent of field theory, I tend to find this sort of symbolic behavior influential, but we just don’t know how influential it is and for what subset(s) of the field (that’s one weekness of field theory: individual actors do have agency and can resist field-level pressures—the question is when and how can they do that).
The most systematic and relevant evidence we have is on research norms, coming from survey and interview data with sociologists. These studies suggest deep contestation in the field over the standards of research production, especially over topics like whether objectivity is not just possible but even desirable, whether scholars should separate advocacy or activism from their research, and whether it is acceptable for scholars to ask certain research questions or evaluate certain explanations (Akresh 2017; Gross 2013; Horowitz et al. 2018). These data are now out of date and I would love for someone to update them, since I suspect we’ll see movement just based on high-profile commentaries by big-name scholars at good places as well as the willingness of a lot of folks to publicly say things I don’t think a lot of folks would have said publicly just ten years ago.
However, there are fewer systematic studies evaluating the quality of research produced in the field.
Again, the best evidence on this front comes from more superficial measures: the research produced tends to skew, and has become more skewed, toward topics related to power and inequality, especially relating to race, gender, and sexuality, specificially the traditional minorities of each category (Altman and Cohen 2025; Heiberger et al. 2021; Keskintürk 2023; Moody and Light 2006), while overlooking other topics central to the systematic study of society but that have become either taboo (off limits) or are simply blindspots (it doesn’t occur to most to study these, there’s little interest in studying them, or there’s little demand/incentives in funding or hires for them). For example, scholars have pointed out sociologists’ reticence to study biology and evolution (e.g., Bearman 2008; Heuveline 2004; Horowitz et al. 2014; Whitmeyer and Hopcroft 2025); religion, especially beyond a narrow focus on “Christian nationalism” (Heiberger et al. 2021; Pearce 2023; Smith 2024); and some topics like socialization (Guhin et al. 2021) associated with allegedly conservative values of order and stability. Similarly, scholars have pointed to explanations and theoretical frameworks considered taboo, including sociology’s shunning of cultural explanations for racial disparities (Horowitz et al. 2018; Patterson 2014); a focus on racism instead of classism in explaining racial disparities (Horowitz et al. 2018; Iceland and Silver 2024; Wilson 2011); and the rejection of structural-functionalist theoretical frameworks as conservative (Iceland 2025). Scholars have also described a negativity bias in sociology to some extent traceable to the field’s focus on “social problems,” but to some extent to ignoring (likely unintentionally) contrary evidence of improvement or simply failing to sustain an applied wing or explore the conditions under which certain positive outcomes occur (e.g., Best 2001; Iceland 2025; Martin 2016; Rojas 2025).
However, beyond pointing to high-profile examples or collecting examples from across the field of bad work that shouldn’t have been published (or published in its final form) or received certain awards, we do not currently have clear, systematic evaluations of the quality of the research produced. Indeed, we cannot determine, based on available studies, the extent to which efforts like citation justice or decolonizing the field have corrupted the research produced. We do have some systematic reviews of research, however, that cover some important topics. Recent studies have documented the extent to which research across fields is not politically neutral (Manzi 2026) and fails to replicate or reproduce (but it is not clear whether these problems are related to the authors’ political biases or not). However, a common question remains whether those and other quality issues are inversely correlated with journal prestige: that is, are these documented problems field-level problems, but not problems in the top journals?
The most difficult area to explore systematically involves the formal and informal sanctioning of scholars for producing taboo work, precisely because such sanctioning is rarely publicly documented. Thus, while it is not the case that scholars are never able to publish controversial work or have successful careers after doing so, several high-profile cases have been documented of scholars being ostracized for their work[1] or fielding challenges to their tenure cases.[2] While some scholars have discussed various forms of self-censoring (LaFree 2025), reviewer criticism using activist standards (Rubin 2025), and retraction (Savolainen 2024), we do not have systematic data on how widespread (if at all) such cases are. Here, survey work would be welcome.
One area where I’d really like to see future research would be investigating how much of a chilling effect high-profile controversies had (and how long they lasted, in the case of fields where those controversies have died down). My own speculation is that these sorts of episodes don’t have to be widespread for folks to self-censor or write in a way that anticipates and tries to stave off political critiques. I suspect the latter itself helps to further shift the norms in the field about what is appropriate and inappropriate to say or write in one’s research.
Another tricky possibility to investigate is that we could have a Keynesian beauty contest going on: people base their behavior on what they think other people are thinking. Folks could do surveys asking folks what they think the majority of scholars (or their friends) think, which is sometimes more effective (accurate) than simple polling on what the surveyed person thinks. The question is whether we are in that scenario—that is, is the majority just adjusting to fit an ideal very few people prefer—or if folks aren’t adjusting, they’re just there already (or if it’s a mix—folks are halfway there, but they keep adjusting because the field keeps moving and they’re trying to keep up). At this point, we’re all embedded in different, nested or overlapping populations; when I ask folks what they think, I get wildly different takes on the answer to this question.
Now for my speculation: my own sense is that some (most??) high-quality scholars continue to have high scholarly standards, but the majority of scholars (even some really good scholars) have let their standards erode (and that’s at least consistent with some survey data, as well as what I saw as an editor for a journal in a cognate field). But as I’ve mentioned before, that’s pretty endogenous—the folks I call good scholars are those who espouse and follow what I view as scholarly standards (except that you do see some formerly “good scholars” who have relented on some of these issues, so there is some drift even at the top). More concretely, I’m fairly certain there is a strong generational divide with academically older (i.e., more years since the PhD) faculty preferring, on average, the high scholarly standards of yesteryear and newer scholars and grad students, on average, preferring activism (political impact), and folks in my generation are more a mix. Again, heterogeneity. And that makes it really difficult to generalize and easy for folks to say, well that’s never happened to me so it must be wrong, while others think what I call scholarly standards are “problematic” or inherently political. [So prescient….]
The Not Clearly Political Issues
Now for the other stuff. There is a longstanding discussion about disciplinary fragmentation/siloing and theoretical stagnation in the field (too many to cite here). I went back and forth about writing about this in my report, and whether I needed to include a COI (which I ultimately did). I think the political stuff did exacerbate these issues, but I’m not sure I can support that with evidence (I’ve also speculated about this in my own published work). But say we do eventually strip away all the political stuff; we’ll still need to deal with this central issue—the field is siloing itself and with siloing comes theoretical stagnation. Obviously (see Heyer et al. 2023, 2025, 2026, Rubin and Shalaby 2025), I think this is a pretty important issue, and it’s not limited to sociology. And as I’ve written before, I do think it’s related to an erosion of scholarly standards—some of which has been justified on moral/political grounds (e.g., inclusion, anti-gatekeeping, a redistribution mindset), but some of it is also just structural or practical (and Alena and I enumerated some of the reasons for this).
I also have a whole section on other issues the field is facing that I mention in passing. I mention a lot of them in this earlier post, so I won’t rehash them here. One I don’t mention in that post, but I do in the report, is sociology’s gender skew issue and refusal to talk about it, which Phil Cohen has written about here and here. There are a few other issues, including some that I added after my report was done. Others have certainly argued that these issues are directly related to the field’s politics, but I’m not convinced by the evidence that they are (which isn’t to say they aren’t, just that I don’t think we have the evidence—also, I suspect these other issues are overdetermined and a monocausal account will be insufficient).
Summary
So, like we say in the Report, things are mixed. Sociology is no exception. There’s enough evidence to say there’s a problem, and that problem seems to be worse in sociology than the humanities we looked at, but how extensive the problem is in sociology, and how it manifests, is more complicated and requires more research. [And godspeed to whoever undertakes it because hell’s coming with you.]
Next Steps
I want to close by talking about what I don’t discuss in my report (and what we don’t discuss in The Report). Solutions.
Solutions require that we have a clear causal model for the problem. But at this stage, we are still trying to outline the problem (which only some of us academics agree is a problem). We’re not yet in a position to say with certainty what caused it in any given field and thus how to fix it. My time as an editor gave me a lot of ideas. My work on the commission gave me some other ideas and fixed some misimpressions I had (I, too, had blamed the humanities for problems in the social sciences; now I don’t—I learned that good humanistic research upholds solid standards and the humanities are much stronger than I thought going into this).
The only policy recommendation I feel like I can make is pretty simple: we have to uphold (encourage, incentivize, teach) scholarly standards. I have lots of thoughts on how and why they came to decline—and why people feel so emboldened to openly reject them.
Folks with other policy recommendations have other causal models in mind. I don’t know if they are right about either the causal models or their recommendations (and, for all I know, they could be right). But I can’t say they are right based on my research, which is to say the published research and commentaries out there.
So my policy rec is the minimum one I can make, but it’s also one that roughly half or 60% (at least—again, outdated numbers) would instantly reject. The empirical question is whether it’s enough that the good ones agree with me.
References
Akresh, I. R. (2017). Departmental and disciplinary divisions in sociology: Responses from departmental executive officers. The American Sociologist, 48(3):541–560.
Altman, M. and Cohen, P. N. (2025). The state of sociology: Ev-idence from dissertation abstracts. Working Paper posted on ArXiv https://osf.io/preprints/socarxiv/a8uyp_v1.
Bearman, P. (2008). Exploring genetics and social structure. American Journal of Sociology, 114(S1):v–x.
Best, J. (2001). Social progress and social problems: Toward a sociology of gloom. The Sociological Quarterly, 42(1):1–12.
Gross, N. (2013). Why are professors liberal and why do conservatives care? Harvard University Press.
Guhin, J., McCrory Calarco, J., and Miller-Idriss, C. (2021). Whatever happened to socialization? Annual Review of Sociology, 47(1):109–129.
Heiberger, R. H., Galvez, S. M.-N., and McFarland, D. A. (2021). Facets of special-ization and its relation to career success: An analysis of U.S. sociology, 1980 to 2015. American Sociological Review, 86(6):1164–1192.
Heuveline, P. (2004). Sociology and biology: Can’t we just be friends? American Journal of Sociology, 109(6):1500–1506.
Horowitz, M., Haynor, A., and Kickham, K. (2018). Sociology’s sacred victims and the politics of knowledge: Moral foundations theory and disciplinary controversies. The American Sociologist, 49(4):459–495.
Horowitz, M., Yaworsky, W., and Kickham, K. (2014). Whither the blank slate? a report on the reception of evolutionary biological ideas among sociological theorists. Sociological Spectrum, 34(6):489–509.
Iceland, J. (2025). Beyond conflict: recovering the sociology of social progress. Theory and Society, pages 1–17.
Iceland, J. and Silver, E. (2024). When all you have is a hammer: how social justice distorts what we know about racial disparities. Theory and Society, 53(5):1073–1092.
Keskintürk, T. (2023). Sociology’s inequality problem. Turgut Keskinturk (Blog) (October 10). https://tkeskinturk.github.io/blog/inequality/.
LaFree, G. (2025). Understanding what violent street crime, globalization, and ice cream have in common. Criminology & Public Policy.
Manzi, J. (2026). The ideological orientation of academic social science research 1960–2024. Theory and Society, 55(2):25.
Martin, C. C. (2016). How ideology has hindered sociological insight. The American Sociologist, 47(1):115–130.
Moody, J. and Light, R. (2006). A view from above: The evolving sociological landscape. The American Sociologist, 37(2):67–86.
Patterson, O. (2014). How sociologists made themselves irrelevant. Chronicle of Higher Education (December 1). https://www.chronicle.com/article/ how-sociologists-made-themselves-irrelevant/.
Pearce, L. D. (2023). One hundred years of religion in social forces. Social Forces, 101(4):1633–1643.
Rojas, F. (2025). Sociology of the good society. Temple of Sociology (Substack) (April 23, 2025). https://templeofsociology.substack.com/p/ sociology-of-the-good-society.
Rubin, A. T. (2025). Normativity is not a replacement for theory. Theory & Society, 54(5):795–850.
Savolainen, J. (2024). Unequal treatment under the flaw: Race, crime & retractions.
Current Psychology, 43(17):16002–16014.
Smith, C. (2014). The sacred project of American sociology. Oxford University Press.
Smith, J. (2024). Old wine in new wineskins: Christian nationalism, authoritarianism, and the problem of essentialism in explanations of religiopolitical conflict. Sociological Forum, 39(4):328–340.
Whitmeyer, J. and Hopcroft, R. L. (2025). What makes a good social science theory, and why the evolutionary model of the actor is one. Theory and Society, pages 1–20.
Wilson, W. J. (2011). Reflections on a sociological career that integrates social science with social policy. Annual Review of Sociology, 37(Volume 37, 2011):1–18.
Young, C. and Cumberworth, E. (2025). Multiverse Analysis: Computational Methods for Robust Results. Analytical Methods for Social Research. Cambridge University Press.
[1] For example, James Coleman’s experience discussed on this webpage and this blogpost.
[2] See the case of Mark Regenerus described in Smith (2014). See also a recent reanalysis of his work in Young and Cumberworth (2025). Note that there was also an effort to retract Regenerus’s study. For some commentary from sociologists on the research and tenure case, see https://scatter.wordpress.com/2012/ 06/23/bad-science-not-about-same-sex-parenting/,https://orgtheory.wordpress.com/2012/07/29/comments-on-regnerus/, and https://familyinequality.wordpress.com/2018/03/06/mark-regnerus-to-be-promoted-to-full-professor-at-ut-austin/.
