Epistemology in Science
The companion piece to this one sets out the formal logic behind scientific reasoning: deduction, induction, abduction, and the hypothetico-deductive method that combines them. It ends by showing where the logic alone stops settling anything: a failed prediction refutes a conjunction of theory and auxiliary assumptions, but nothing in the formalism says which conjunct to blame. Everything from that point on is epistemology: questions about what justifies treating one theory as better warranted than another, whether scientific claims about unobservable entities should be believed literally or merely used as means to an end, and what changes when an entire framework of assumptions gets replaced instead of a single hypothesis.
The theory-ladenness of observation
The hypothetico-deductive process assumes that observation reports (the \( O \) in the schema \( T \land A \rightarrow O \)) form a fixed, theory-independent court of appeal to which theories are answerable. Norwood Russell Hanson, and Kuhn after him in a different way, pressed on a complication: what counts as an observation, and how it is described, already presupposes much of the theoretical framework under test. Two physicists looking at a bubble-chamber photograph see different things: one sees a set of curved lines, the other a Compton-scattered electron and its positron partner, and the second description depends entirely on the theoretical apparatus of quantum electrodynamics already being in place.
If observation reports are shaped by the theoretical commitments already in use, a test of \( T \) is never the theory-free trial the modus tollens schema needs. Theories remain falsifiable: the independence of the evidence has to be argued case by case, typically by showing that the theoretical assumptions shaping the observation report are distinct from \( T \), and independently supported.
Scientific realism versus instrumentalism
Once a theory posits unobservable entities (electrons, quarks, spacetime curvature, wavefunctions), a further question opens up: should successful, well-tested theories be believed as approximately true descriptions of what actually exists, or just accepted as reliable instruments for generating correct predictions, agnostic about what, if anything, the unobservable posits correspond to?
The realist's strongest argument is the "no miracles" argument: it would be a staggering coincidence for a theory positing entirely fictitious entities to generate novel, precise, repeatedly confirmed predictions across domains its inventors never anticipated, unless the posited entities (or something close to them) actually exist and behave in the way that the theory suggests. The instrumentalist's reply is the pessimistic meta-induction: the history of science is a graveyard of once-successful theories (phlogiston, caloric, the luminiferous aether) whose central unobservable posits were later discovered not to exist, so current successful theories can have no special exemption from the same fate. Both arguments extrapolate from a pattern (successful theories tend to be true; successful theories have often turned out false) to a general conclusion, and are therefore inductive.
John Worrall's structural realism (1989) offers a middle position between realism and instrumentalism, motivated by trying to hold onto each argument's strongest point simultaneously. Worrall's proposal is that what survives a change in theory is a theory's mathematical structure, not the full ontology of the old theory (the ether as a substance, say): Fresnel's equations for light's reflection and refraction at an interface were carried over essentially unchanged into Maxwell's electromagnetic theory, even though the ether they were originally derived to describe was abandoned entirely. On this view the pessimistic meta-induction is right that theories' unobservable ontologies do not survive, and the no-miracles argument is right that something about successful theories must be tracking reality.
Kuhn
Thomas Kuhn's The Structure of Scientific Revolutions argued that periods of what he called normal science operate within a paradigm (a shared constellation of theory, exemplary solved problems, instrumentation, and standards of what counts as a legitimate question), and that anomalies accumulate for a long time without dislodging the paradigm, because normal science treats anomalies as mishaps to be solved within the existing framework rather than as refutations. This already modifies the naive falsificationist model: a single \( \neg O \) rarely kills a paradigm outright, since the Duhem–Quine thesis gives a working scientist endless auxiliary assumptions to adjust first.
The more contested claim is incommensurability: that during a genuine paradigm shift, key terms change meaning across the transition in a way that makes direct comparison of the old and new theories' claims harder than it looks (the Newtonian "mass" and the relativistic "mass" are not the same quantity with a corrected value, on this interpretation, because the concept's relations to space, time, and velocity have themselves changed). Taken to its strongest form, this threatens the idea of straightforward rational comparison between paradigms and has been read (more often by critics than by Kuhn himself, who pushed back on the strongest readings) as a slide toward relativism about scientific truth. A more modest reading, which most working philosophers of science now take, is that incommensurability is a real phenomenon requiring careful translation and interpretation across a paradigm shift.
Lakatos
Imre Lakatos tried to keep Popper's insistence on genuine testability while accommodating Kuhn's observation that theories are not abandoned after a single refutation. His unit of appraisal is the research programme: a hard core of assumptions held immune from direct refutation by convention, surrounded by a protective belt of auxiliary hypotheses that absorb the impact of anomalies and can be revised without touching the core. A research programme is progressive if successive revisions to the protective belt predict novel facts that are subsequently confirmed; it is degenerating if the revisions are made purely to accommodate already-known anomalies, with no new predictive content: ad hoc rescues in the pejorative sense.
This gives a principled answer to the question the Duhem–Quine thesis leaves open in the companion piece: when a prediction fails, which conjunct should be blamed, the theory or an auxiliary assumption? Lakatos's answer is that the hard core should be protected as a matter of research strategy, provided doing so continues to generate novel, confirmed predictions; the moment the programme can only explain what is already known, without predicting anything new, continued core-protection becomes an epistemic and academic vice.
A worked case: continental drift, from degenerating to progressive
Wegener's 1912 hypothesis, that the continents had once formed a single landmass, Pangaea, and had since drifted apart, is an unusually clean illustration of the Kuhn/Lakatos machinery operating on a real, contested theory over an extended period, unlike the idealised single-test snapshot the hypothetico-deductive schema suggests. Wegener's evidence for \( T \) was abductive: the jigsaw fit of the South American and African coastlines, matching rock strata and fossil distributions (the reptile Mesosaurus and the fern Glossopteris) on both sides of the Atlantic, and matching mountain-belt ages across continents now separated by ocean. Each of these facts was, individually, compatible with rival explanations already accepted within the geological paradigm of the time (land bridges since submerged, or independent parallel evolution), and the geological establishment, overwhelmingly, preferred those explanations to Wegener's.
The decisive objection was an auxiliary-assumption failure in the Duhem–Quine sense. Wegener could not provide a physical mechanism by which continents, made of less dense granitic rock, could plough through the denser basaltic ocean floor, and the mechanisms he did propose (tidal and centrifugal forces) were shown by physicists, including Harold Jeffreys, to be many orders of magnitude too weak. On Lakatos's terms, continental drift in its 1912–1920s form was a degenerating research programme with respect to mechanism: it accommodated existing evidence but could not generate a physically credible novel prediction about how drift occurred, and defenders had no better response than to set the mechanism problem aside and hope. Geologists at the time were applying a criterion (demand a mechanism, or at least demonstrate that the proposed one is physically adequate) that is a perfectly reasonable epistemological standard, applied consistently elsewhere in physical science; downgrading continental drift on that basis was not irrationality or conservatism.
What changed the verdict, decades later, was a new observational domain supplying the missing mechanism: palaeomagnetic striping on the ocean floor, discovered in the 1950s and 1960s, showing symmetric bands of alternating magnetic polarity mirrored on either side of the mid-ocean ridges, exactly as predicted by seafloor spreading driven by mantle convection. This was a novel prediction in Lakatos's sense (nobody was looking for magnetic striping in order to rescue Wegener; the striping was predicted by the seafloor-spreading mechanism and then found), and it converted the programme from degenerating to sharply progressive almost overnight, with plate tectonics as the resulting hard core absorbed into geology within a single decade. The case is instructive as both verdicts were reasonable given what was known at the time: rejecting mechanism-less continental drift in 1920 and accepting mechanism-backed plate tectonics in 1965 reflect a single research programme's status changing, as Lakatos's criterion predicts, when its predictive fortunes changed.
Feyerabend
Paul Feyerabend's Against Method (1975) took the historical case of Galileo's arguments for Copernicanism and argued that no single methodological rule (falsifiability, Lakatos's progressive/degenerating criterion, or any other proposed demarcation standard) survives the history of science without being violated. Galileo, on Feyerabend's reading, defended heliocentrism with arguments that were, by the observational and methodological standards of his own time, no stronger than the Aristotelian alternative, and progress came through a principled rule-breaking rather than rule-following.
The more defensible core of the argument is methodological pluralism: a fixed, universal algorithm for good scientific method, applied rigidly across every historical episode, would have blocked some of the theory changes now regarded as scientific genius. The criteria this piece has surveyed (falsifiability, novel prediction, progressive-versus-degenerating status) function better as retrospectively applied standards of appraisal than a priori rules a scientist should consciously follow while doing research.
What separates a scientific explanation from a valid prediction
Everything so far has treated \( (T \land A) \rightarrow O \) as a process for prediction and testing, but the same schema is offered as a model of explanation as well: Hempel and Paul Oppenheim's 1948 deductive-nomological (D-N) model holds that to explain an event \( E \) is to deduce it from a set of general laws \( L \) together with statements of antecedent conditions \( C \), \( L \land C \vdash E \), the same logical form as a scientific prediction, differing only in whether \( E \) is already known (explanation) or not yet known (prediction). This "symmetry thesis," that explanation and prediction share one logical structure, is elegant and looks like it should be the end of the matter.
Two classic counterexamples show that valid D-N form is not sufficient for explanation. The asymmetry problem: the height of a flagpole, together with the law of rectilinear light propagation and the sun's angle of elevation, deductively yields the length of its shadow, and this is a perfectly good explanation of why the shadow has the length it does. But the shadow's length, the same law, and the sun's angle equally deductively yield the flagpole's height, in an argument of identical logical form. Yet nobody accepts "the shadow explains the flagpole's height," because the flagpole causes the shadow. The D-N schema, being symmetric in \( \vdash \), cannot distinguish the two arguments; something about explanation that is not captured by deducibility is implying a symmetry break. The irrelevance problem, due to Wesley Salmon: "all men who take birth control pills fail to get pregnant; John takes birth control pills; therefore John fails to get pregnant" is a valid D-N argument from a true general law and a true antecedent condition to a true conclusion, and it explains nothing, since John's not getting pregnant has nothing to do with the pills and everything to do with being male.
Wesley Salmon's alternative, the causal-mechanical model, responds by moving the explanation's core requirement away from deducibility and into the structure so that to explain \( E \) is to trace the processes and interactions that produced it, which handles the flagpole case (light travels from sun to flagpole to ground) and the birth-control case equally as elegantly (the pills are causally inert with respect to John's sex, whatever law-like generalisation can be stated about the pill user). Philip Kitcher's rival unificationist account keeps explanation tied to derivation, as Hempel wanted, but adds a global constraint absent from the original D-N schema: genuine explanations are instances of the smallest set of argument patterns capable of deriving the largest range of accepted beliefs, so a good explanation unifies otherwise diffuse phenomena under a shared derivational pattern rather than only satisfying \( L \land C \vdash E \) for some \( L \) and \( C \) found after. Both replies agree that Hempel's insight survives, that explanation has a lawlike, general structure rather than being a matter of narrative or psychological satisfaction, and both agree that pure deducibility, policed by form alone, is not sufficient or useful.
Values in science and the problem of inductive risk
A separate argument, associated with Richard Rudner's 1953 paper and developed further by Heather Douglas and others since, holds that non-epistemic values are an unavoidable component of scientific reasoning, given how the hypothetico-deductive process gets applied. Accepting or rejecting a hypothesis on the basis of evidence always involves setting a threshold of evidential strength required before acceptance, and how demanding that threshold should be depends on the costs of being wrong in each direction: this is the problem of inductive risk. Setting an unsafe threshold for evidence that a new drug is not harmful trades a lower false-negative rate (missing a real harm) for faster access to a possibly beneficial treatment and setting a stricter threshold trades the reverse.
This does not license "science is politics," which is a conflation Douglas is explicit about resisting. Helen Longino's related feminist philosophy of science makes a complementary point: background assumptions bridging data to hypothesis (which auxiliary \( A \) to hold fixed) are frequently missed because they are widely shared within a scientific community, and a community with a narrower range of social perspectives is correspondingly less likely to notice when a background assumption is doing unacknowledged work. Longino's proposed remedy is to cultivate the kind of transformative critical discourse (open venues for dissent, uptake of criticism, publicly shared standards, equality of intellectual authority among participants) that gives a community the best chance of surfacing and correcting assumptions that would otherwise be missed.
Social epistemology: peer review, replication, and trust
Individual scientists cannot personally re-derive, from first principles, more than a small fraction of the background knowledge any single piece of research depends on. Scientific knowledge is, unavoidably, held collectively and transmitted substantially by testimony. That raises its own epistemological questions such as when is it rational to accept a claim on another scientist's authority rather than independently verifying it? Peer review functions as an attempt to answer this at scale: a filter that catches a subset of methodological errors before a claim enters the literature and becomes something other scientists build on and cite.
The replication crisis, most visibly documented in psychology and biomedical science from around 2011 onward (the Open Science Collaboration's 2015 attempt to replicate 100 psychology studies succeeded, on the strictest criterion, for only about 36 of them), is best understood as an empirical discovery about how well this institutional filter was working, rather than a discovery about any single theory being false in the modus-tollens sense. The mechanisms identified are all failures at the level of the community's incentive structure rather than failures of any individual scientist's logic: publication bias favouring positive, novel results over null ones; \( p \)-hacking (analysing data multiple ways until a result crosses the conventional \( p < 0.05 \) significance threshold, then reporting only that analysis); underpowered sample sizes; and the historical absence of institutional reward for the unglamorous work of direct replication. This connects directly back to inductive risk: a discipline-wide norm of accepting \( p < 0.05 \) as "significant," without separately asking whether that threshold suits the costs of false positives in a given field, is a case of the value-laden threshold-setting described above, made invisible by being a fixed convention rather than an explicit judgment argued case by case. The corrective measures now widely adopted include pre-registration of analysis plans before data collection, mandatory replication attempts, and larger, better-powered samples. Together they amount to institutional epistemology: engineering the community-level structure so that the collective output tracks truth better than the previous structure did, a problem no amount of individually correct reasoning could solve from the inside.
The file-drawer problem, closely related, conveys the same point in a form that connects directly to the Bayesian tools from the companion piece. If ten independent laboratories each test a hypothesis at the conventional 5% significance threshold and the null hypothesis happens to be true, ordinary probability says that roughly one of the ten will nonetheless produce a "significant" false positive by chance. If only that one laboratory publishes, since null results are harder to place in high-profile journals and less career-advancing to write up, the published literature contains a single confident positive finding with no trace of the nine null results that would have correctly contextualised it. No scientist in this scenario reasoned invalidly, and no experiment was conducted improperly: the distortion is a property of the publication process. This is why meta-analysis and pre-registered study registries (recording a trial's existence and design before its results are known, so a null result cannot vanish from the record) have become standard correctives. It demonstrates that a community composed entirely of individually rational, logically careful scientists can still produce a collectively misleading body of literature, purely as a function of how results are recorded.
Laudan
Larry Laudan's Progress and Its Problems (1977) offered a further reconstruction of scientific change, developed as a response to a weakness he saw in both Kuhn and Lakatos: both still tie a paradigm's or programme's success, ultimately, to some notion of approaching the truth or of novel predictions being confirmed. Laudan argued that scientific progress can be tracked, and rival theories compared, using a criterion that does not presuppose realism about unobservables. On Laudan's account, a theory's job is to solve problems, of two types: empirical problems, which are the phenomena a theory is called on to explain or predict (Hempel and Popper's territory), and conceptual problems, internal or external tensions a theory generates regardless of how well it matches observation: internal conceptual problems are inconsistencies or ad hoc patches within the theory itself, external ones are a clash between the theory and some other well-supported theory or deeply held methodological principle.
Discussion