Since we are told all the time that mathematics and measurement and evidence are the foundation of science, some of you may not believe me when I tell you that there is no accepted or widely recognized mathematical measure of evidence within science. Science routinely deals with math and measures and evidence, but it seldom uses math to measure evidence.
Everyone recognizes the importance of evidence. But unlike most people, I believe it to be the single most important concept we need to make sense of the world. I use ‘evidence’ as a synonym for ‘knowledge’ and ‘information,’ but it is a superior term insofar as it emphasizes uncertainty, it is more explicitly relational, and it is more precise since it has a polarity (information about hypothesis H can refer to evidence for or against H).
I ultimately want to understand and measure the evidence in a cognophysical model of a thing/observer. However, that is a strange and unfamiliar notion. It is easier to first understand our own evidence about the world, as in the case that we have scientific observations that are evidence for some external reality.
For example, this could be a photo of galaxies taken with the aid of a telescope. We say that this is “a photo of galaxies,” but that is only a hypothesis. The photo is a two-dimensional pattern of light that we observe on a screen or paper. It provides us with “data” that we interpret as evidence for the hypothesis that real galaxies having particular properties exist, and they caused the observed data. How strong is the evidence in data D for hypothesis H?
Science Measures Prediction Error, Not Evidence
Scientists seldom measure evidence. What they routinely measure is prediction error, which is inversely related to the accuracy of prediction. The more accurately model/hypothesis H predicts data D, the stronger the evidence D for model H. But how much stronger?
A measure of predictive accuracy, such as mean squared error, is not a measure of evidence. This can be seen in its units, which could be meters or kilograms. There is no simple or general relation between prediction error and evidence. One cannot calculate an amount of evidence given only the magnitude of a prediction error.
A science that relies on measures of prediction error to select its models, rather than measures of evidence, naturally produces good predictive models (algorithms) rather than good models of reality (see Science’s Problem with Reality). A good model of reality must make accurate predictions, and thus it must be a good predictive model. But a good predictive model need not be a good model of reality (as in the example of epicycles; see below, and Science’s Problem with Reality).
Science without a means of measuring evidence is like geometry without a means of measuring distance. Science needs a measure of evidence that is just as rigorous and widely understood as the measure of distance in geometry.
Probability Objectively Measures Evidence
I say that probability is an objective measure of evidence, so that probability P(H|D) measures the evidence in data D for hypothesis H. I recently introduced axioms that formalize probability as a measure of evidence (Fiorillo et al., 2025). But the majority of scientists have not taken this view, even though they do not endorse any alternative measure of evidence. Among those who do believe that probabilities measure evidence, at least in principle, there is no consensus about the exact and general relation of probability to evidence (how to assign a probability given evidence). As a result, the probabilities that scientists derive often do not measure what they are said to measure, or they do not measure it correctly.
I believe that Harold Jeffreys has done more than anyone else to establish a measure of evidence, in what is now known as “Jeffreys prior” (though he did not refer to it as a measure of evidence). I have been working to better justify his measure and to relate it to propositional logic and the probability axioms of Cox and Kolmogorov. Here I provide only a general and crude background to this work. More detail can be found in Kim and Fiorillo (2016) and Fiorillo and Kim (2021), and especially Fiorillo et al. (2025). I give an example further below of measuring evidence in a simple geometric model of an observer.
Probability Theory: The Logic of Science
I read Probability Theory: The Logic of Science by the physicist E.T. Jaynes while I was a postdoctoral student in 2003. It had a much greater influence on me than anything else I ever read, and radically changed the course of my research and my career.
It was not the math itself that had a great influence on me. Laplace referred to probability theory as “nothing but common sense reduced to calculation.” Common sense is reason, and reason is mathematics, and there is nothing particularly interesting about it. I did not even pay much attention to the math when I first read the book.

The big idea addressed by Jaynes concerns epistemology, the nature of knowledge and how we can measure it. After reading it I felt that I understood the purpose of the brain in a quantitative sense (even though I was not yet able to assign numbers). Today I express this most concisely by saying that the purpose of the brain is to maximize its evidence for its future existence.
The general form of a probability is P(A|B). This should be understood to measure the evidence in proposition B for the truth of proposition A. The easy problem in probability theory is to calculate probabilities from other probabilities, which is merely application of established and uncontroversial mathematics. The hard problem is to specify the relationship between the evidence within propositions (A|B) and the numerical measure of that evidence P(A|B) (in other words, to derive probabilities from information, so as to assign numbers to P). There are still no methods of doing this that are both universally applicable and widely accepted. I have been working toward a system of axioms that provide an objectively correct measure of evidence.
Conceptual Obstacles and Alternative Definitions of Probability
There are a remarkable number of conceptual obstacles concerning probability theory that have hindered progress. I summarize a few of these below. Jaynes (2003) addresses these at length. I have addressed them as well (Fiorillo, 2012; Fiorillo and Kim, 2021). I recently introduced axioms that should help to overcome some of the conceptual obstacles by resolving technical shortcomings in the standard axioms of Kolmogorov and Cox (Fiorillo et al., 2025).
Objectivity and Subjectivity
There has been confusion about objectivity and subjectivity. Objectivity is obviously desirable. At the same time, probability is clearly related to beliefs, and those are subjective. We want probability to be able to deal with the reality that different observers have different beliefs (which I equate with evidence). It may seem that we need two types of probability, one objective and the other subjective.
We do not need two types of probability. ‘Objective’ should be understood only to mean rational/logical. Probability theory should be just like other branches of pure mathematics, and pure mathematics is pure logic.
We also want probability to be able to deal with the reality that different observers have different evidence, and this makes evidence subjective. There is no conflict here, because “objective” refers to reason/logic, which is universal and timeless, and “subjective” refers to the fact that evidence is local in space and time, and thus varies from one observer to another. All evidence is subjective in this sense, but that does not make it any less real or valid or rational. When properly understood and applied, probabilities are objectively correct measures of real and subjective evidence.
Note that in common usage, we call a statement “objective” when there is a consensus among well informed observers that it is true. This naturally occurs when observers share the same evidence. But this is merely a special case of the general rule that evidence tends to vary from one observer to another, and thus tends to be local and subjective. To refer to this special case as “objective” causes needless confusion, at least in the context of probability theory.
Frequentist Probabilities
One attempt at objective probabilities has been the “frequentist” definition of probability, which says that probability measures frequency. This has been a prominent misconception. Its flaws have been well documented by Jaynes (2003) and many other “Bayesians.”
If we flip a coin an infinite number of times, and we observe heads in 43% of cases, then frequentists will say that the probability of heads is 43%, regardless of what else we may know about the coin.
This is useful, especially if you have a lot of time to flip coins before you place your bet.
But it is not very useful. For example, if you have not observed the outcome of any coin tosses, and all you know is that a coin has two sides, heads and tails, frequentist probability simply does not apply. It measures nothing, because although there is evidence (we know there are two sides), there is no frequency. Even if you are an exceptional coin tosser and you spend a week of your vacation observing the outcomes of a few trillion coin tosses, you still cannot be certain of the “true frequentist probability.” That would require an infinitely greater number of observations.
In contrast, Bayesians calculate exact probabilities, and have a lot more free time. They simply note that if all we know is that the coin has two sides, the evidence for each must be equal, and therefore the probability of heads is exactly 1/2. There is no need to observe anything in order to mathematically measure evidence.
Objectivist Bayesian Probabilities
Objectivist Bayesians see probability as a measure of “rational degrees of belief.” I share this view. However, I see rationality as inherent to nature, and I see no useful distinction between “belief” and “evidence.” Indeed, I believe the physical state of a brain, or any other thing, to be evidence for its own future existence, as well as evidence about the external state of its environment. I therefore use ‘belief’ and ‘evidence’ interchangeably, though I strongly prefer ‘evidence.’
Objectivist Bayesians include Laplace, Jeffreys, Cox, and Jaynes. Objectivist Bayesians maintain that if A and B are sufficiently well defined propositions (hypotheses), then there is a number P(A|B) that is a uniquely correct measure of evidence B for A. I think everyone agrees this to be desirable. The problem has been the absence of a universally applicable method to find the correct number P(A|B). I have called this “the measurement problem” of probability theory (Fiorillo et al., 2025).
To address this problem, Jeffreys applied principles of invariance to justify what is now known as Jeffreys prior. Jaynes, following the work of Shannon on “information theory,” introduced his principle of maxiumum entropy. Though I believe these methods to be valid, their validity is not immediately obvious, and they have not been universally understood or accepted.
What is obvious is that if B provides no evidence favoring the truth over the falsehood of A, then the evidence for each must be equal. This is known as the “principle of indifference” and results in uniform probabilities. It has always been the foundation for probability theory. However, although everyone accepts it for certain cases, I am not aware of anyone who has believed it to provide a general solution to the measurement problem. Certainly Jaynes did not believe it to be sufficient, despite his vigorous defense of its validity.
My colleagues and I have recently provided what I believe to be a proof that uniformity is necessary and sufficient to derive every probability (Fiorillo et al., 2025). Indeed, uniformity is pure logic and it underlies every use of numbers to measure anything. The central challenge is only to specify exactly what is possible given some model of reality (i.e., to specify a complete set of indivisible propositions).
We cannot actually know what is possible in reality, and thus we need a model (a set of propositions). But I believe it to be obvious that, given what we do know, the set of possibilities is infinite. Therefore we need infinite numbers. We also need infinite numbers that differ in size. For this purpose our axioms use “hypernatural numbers.” They allow for uniform and infinitesimal probabilities over infinite sets, thereby overcoming the well known problem of “improper” (unnormalizable) probabilities that has plagued the axioms of both Kolmogorov and Cox.
Subjectivist Bayesian Probabilities
There are subjectivist Bayesians who do not accept the objectivist ideal. They include de Finetti, Berger, and Bernardo (all accomplished mathematicians at the less subjective end of a subjectivist spectrum). They agree with objective Bayesians that evidence is subjective insofar as it is local in space and time, and tends to differ from one observer to another. But unlike the objectivists, they are skeptical that it is even possible for probability to be an objectively correct measure of evidence.
I see the subjectivist perspective as the result of trying to assign probabilities to propositions that are not sufficiently well defined. We inevitably have ambiguous probabilities when we begin with ambiguous propositions. It is no coincidence that subjectivist Bayesians have tended to work in the social sciences (de Finetti was an economist), where the tolerance for ambiguity is necessarily greater than physical sciences. Objective Bayesians have tended to be physicists, including Laplace, Jeffreys, Cox, and Jaynes.
The Evidence in a Simple Geometric Model of an Observer
We would like to measure the evidence within a simple but realistic physical model of an observer. This means quantifying what one physical thing knows about another. In more standard and technical terms, we partition a physical system into “internal” and “external” portions (or present and future). This means that we split our model, in our minds, into these two portions for purposes of calculation. Then we find the probabilities over possible external states conditional only on the known internal state. This internal state is “an observation.”
I equate “observation” with “observer.” This is because my interest is an observer, and evidence, that is local in space and time. Therefore there can be no distinction between observer and observed and observation. In common language I say that I am the same person and observer over time, but this is not precise. My conscious perception changes multiple times per second. Every time it changes, I become a new observer/observation, even if the change is small.
The physical model my students and I have considered is simply a configuration of geometric points (or “point-particles,” but without mass or velocity or any properties other than relative location). The simplest case to consider can be called “the triangle problem.” There are 3 points, and our observer is 2 of these points. Given only knowledge of the location of 2 points (and thus the distance between them), where is the third point?
We cannot deduce its location, but we can infer it from nothing but the two known locations. It might seem that the best we can do is to say that the third point could be anywhere, and all possible locations are equally likely. However, the problem is much more nuanced than that. Before we can even consider whether all locations are equally probable, we must determine what is possible. What are the “wheres” in “anywhere?”
What is Possible?
We do not know what is possible. To deal with any problem in physics, or in life, we first want to know what is possible. Neither geometry nor physics has a compelling answer to a question as simple as where this third point could be. It is not as though we could do an experiment to find the answer. Possibilities are not observable. Reason is a much better guide here than experience.
There was a famous argument related to this topic between Isaac Newton and Gottfried Leibniz (known as “the Leibniz-Clarke Correspondence”). Newton’s model of mechanics assumes that there is an infinite three-dimensional space that exists and is always the same, regardless of what, if anything, is in it. This is known as “absolute space.” If Newton is right and this space really does exists, then the third point in our triangle could be anywhere in this space.
Leibniz rejected the hypothesis of absolute space, at least in part because it is entirely unobservable. What we can observe is only relations between objects, such as the ratio of the lengths of two sides of a triangle. Leibniz argued that there exists only relations between objects, which implies that there is only a relational space of possibilities. Ernst Mach, and recently Julian Barbour, have built upon this idea.
It seems most physicists believe Newton was right. They may be sympathetic to the argument of Leibniz, since we would certainly prefer not to assume the existence of unobserved entities in our models, let alone unobservable entities. But the model of Newton (as modified by Einstein) has been highly successful in predicting our observations. Furthermore, neither Leibniz nor Mach proposed a specific relational space, let alone a specific model of how things move through it. Since the primary concern of most physicists is accuracy of prediction, and the absolute space of Newton and Einstein works well in that regard, most physicists accept it (some may think Einstein’s model does not include absolute space, but that is incorrect).
There is no substantial evidence for Newton’s absolute space. As discussed above and in Science’s Problem with Reality, predictive accuracy is not the same as evidence in a mathematical sense. Furthermore there are countless conceivable models of reality that all make the same predictions of what we will observe. There are in fact an infinite set of models, of space and “laws of motion,” that would all predict the same observations that Einstein’s model predicts.
The Minimal Model
When we make a model, we are like God creating the universe, and we can specify whatever knowledge and possibilities we like. The space of possibilities in our model could be anything we choose, and there are an infinite number of conceivable spaces. Why should we choose one over another?
There is in fact a minimal model of space. It is the model that assumes the least. In the present case it is the space that is logically implied by the 3 geometric points alone, without any additional assumptions. It can be proven that this space really does exist, in the imaginary sense in which mathematicians use “exist.” I could make a more persuasive case here were I ready to actually specify what this space is, but practical considerations do not permit that yet.
It can also be proven that this model and its space are preferred by reason alone over all other models with 3 points. “Reason alone” means in the absence of empirical evidence, or before any empirical evidence is considered. Therefore this is the best model, until empirical evidence indicates otherwise.
Before finding the probability that the third point is in a particular location, we must first choose a set of parameters that can specify a triangle, such as the lengths of its three sides (a,b,c). If we designate our observer/observation to be the two points that constitute side a, then our observer knows length a, and we want to find P(b,c | a). We will have one probability, determined by reason alone, for each set of propositions. For example, we could find the probability for the case that the propositions (a,b,c) are 0.95 < a < 1.05, 0.80 < b < 0.90, and 1.10 < c < 1.20.
Our observer here is only 2 points, and is simpler and more ignorant than any such observer we can imagine (except 0 or 1 point). But even without knowing the solution or how to find it, it should be apparent that this observer knows quite a lot. This is because there is a highly informative relation between the side lengths of a triangle. Indeed, there are few possibilities here relative to a model that features Newton’s absolute space (which is infinite in each of 3 dimensions).
Rational Observer Theory will show that this minimal model of an observer is more interesting than it may appear. It provides at least a first step towards understanding the physical basis of evidence/observation.
