From Observation to Hypothesis
You notice something. A handful of charts show price stalling at a round number. A stock you follow seems to fall on Mondays. Several of the biggest winners in your watch list made a new 52-week high before they ran. Each of those is an observation, and an observation is not yet anything you can test — it is a report of what your attention landed on.
This lesson is about the conversion. By the end you should be able to take a sentence beginning “I noticed that” and produce from it a written claim that has a population, a condition, an outcome, a horizon, a comparison, and a stated way of turning out to be false. You should also be able to look at a claim somebody else wrote — including your own from last month — and say precisely which of those six pieces is missing.
An observation is a report about your attention
Section titled “An observation is a report about your attention”The trouble with an observation is not that it is wrong. It is that it carries no information about how it was produced. When you say “price often stalls at round numbers”, you are reporting the output of a process with several stages you did not control:
- which charts you happened to open, and why you opened those ones;
- how many bars you scrolled past without noticing anything;
- what counts, to your eye, as “stalling”;
- the fact that stalls near round numbers are memorable and stalls near 47.31 are not.
None of that is a criticism of you. It is how attention works, and it is also how most good ideas start. But the same process reliably produces observations about things that are not there, and it produces them with exactly the same feeling of confidence. A hypothesis is the device that lets you tell the two cases apart.
Where this lesson sits
- ObservationSomething you noticed
- HypothesisA claim that could be false
- RulesNo room for interpretation
- AFLSignals a machine produces
- EvidenceRead against the claim
The six parts of a testable claim
Section titled “The six parts of a testable claim”A hypothesis a computer can evaluate contains six things. Miss any one and the test you eventually run will answer a question adjacent to the one you meant.
| Part | The question it answers | A weak version | A usable version |
|---|---|---|---|
| Population | Which instruments, over what period? | “Stocks” | “Shares in this watch list, 2005 to 2019, with 50-day average turnover above 5,000,000” |
| Condition | What has to be true for the claim to apply? | “After a strong run” | “Close above its own 200-day simple moving average on the previous bar” |
| Outcome | What are you measuring? | “It goes up” | “Return from the next open to the open 20 bars later, after 0.1% per side” |
| Horizon | Measured over how long? | “Soon” | “20 trading bars” |
| Comparison | Compared with what? | (absent) | “The same measurement on all bars of the same symbols in the same period” |
| Falsifier | What result would count against it? | (absent) | “A mean difference at or below zero, or a difference present in only one of three sub-periods” |
The comparison and the falsifier are the two that get left out, and they are the two that do the work. Without a comparison you cannot tell a real effect from the general tendency of the market during your test window: a rule that returns eight per cent a year in a period when holding everything returned twelve is not a discovery, it is a subtraction. Without a falsifier written in advance, no result can disappoint you, because you will find a reading of any number that leaves the idea intact.
Falsifiability, in a trading context
Section titled “Falsifiability, in a trading context”Karl Popper’s criterion translates awkwardly into markets, because nothing here is settled by a single experiment. A weather forecast can be checked tomorrow; a claim about a 20-day edge in a universe of 500 shares is checked against a distribution, and distributions do not falsify things cleanly. What survives the translation is the practical half of the idea: a claim is worth testing to the degree that you can say in advance what evidence would make you drop it.
Three patterns make a trading claim unfalsifiable in practice. All three feel like sophistication rather than evasion.
The conditioning escape. “It works in trending markets.” Fine — but if “trending” is identified after the fact by whether the rule worked, the claim has no content. The fix is brutal and simple: define the regime with a rule that could be evaluated on the morning of the trade, using only data available then, and make it part of the hypothesis rather than a commentary on the result.
The discretion escape. “Take the signal, unless the chart looks wrong.” Every backtest of such a rule tests a different system from the one you would trade, and no result transfers. This is not an argument against discretion in trading; it is an argument that a discretionary overlay cannot be evaluated by a machine, so it must be excluded from the claim and assessed separately, as Part 35 does.
The moving target. You test 20 days, get nothing, try 15, then 25, then add a volume filter, then restrict to one sector, and stop when a number pleases you. Each step is reasonable in isolation. Together they are a search, and the end of a search is not evidence about the thing you started with. Part 30 measures how badly this inflates results; the defence is to write the hypothesis down first and to record every subsequent variation as a separate, later, weaker test.
Why a mechanism matters
Section titled “Why a mechanism matters”You can test a claim with no mechanism behind it. Machines will happily compute the average 20-day return of shares whose ticker contains the letter Q. What a mechanism buys you is protection against a specific, extremely common failure: finding a pattern because you looked at many patterns.
A mechanism is a sentence about why the pattern would exist, phrased in terms of somebody doing something. Useful families include:
- Risk transfer. Somebody is being paid to hold an exposure others want to shed. Applies to volatility, carry and liquidity provision.
- Information diffusion. News reaches participants unevenly, so prices adjust over days rather than instantly. The standard story behind medium-horizon momentum.
- Structural constraint. A rule forces someone to trade regardless of price — index rebalancing, fund mandates, margin calls, tax-year deadlines.
- Behaviour. Documented, repeatable human tendencies: reluctance to realise losses, anchoring on purchase price, over-extrapolation of recent returns.
- None I can name. A legitimate entry, provided you write it down. It should lower your prior and raise the amount of out-of-sample evidence you will demand.
The practical value is that a mechanism makes side predictions. If your story is that information diffuses slowly, the effect should be stronger in less-covered shares and weaker in the most heavily traded ones. If your story is index rebalancing, the effect should cluster on known rebalancing dates and vanish elsewhere. Those side predictions can be tested, and they are far harder to satisfy by accident than the main claim. An idea that passes its main test and fails every side prediction its own mechanism implies has told you something important.
Write it down before you touch the data
Section titled “Write it down before you touch the data”The single highest-value habit in this part costs about four minutes. Before running anything, write a short, dated record of the claim. Not in your head, and not in a comment you will edit later — in a file whose earlier versions you keep.
A workable format, which the project at the end of this part uses:
- Date and the observation that prompted it, in one sentence.
- Claim, containing all six parts from the table above.
- Mechanism, or an explicit “none I can name”.
- Prior: how likely you think this is, before evidence. A number between 0 and 1 is fine and slightly embarrassing, which is the point.
- Falsifier: the result that would end this line of work.
- Reserved data: the period or the symbols you will not look at until the end.
- Cost assumption: the round-trip cost the claim must beat, decided now.
Keeping the earlier versions matters more than the format. The value is not in the document; it is in being able to see, three weeks later, that the claim you are now testing is the fourth revision of the one you started with, and that the reserved period has been quietly consumed.
Good and bad hypotheses, side by side
Section titled “Good and bad hypotheses, side by side”| Not yet testable | Why | A testable rewrite |
|---|---|---|
| “Breakouts work.” | No population, no definition of breakout, no outcome, no comparison. | “Among shares in this list with 50-day turnover above 5,000,000, a close above the highest close of the prior 50 bars is followed by a mean 20-bar return, measured open to open, higher than the mean 20-bar return of all bars of the same shares in the same period, by more than 0.4% (twice the assumed round-trip cost).” |
| “RSI below 30 means the stock is oversold and due to bounce.” | “Oversold” and “due” restate the condition rather than predicting anything measurable. | “…a 14-period RSI below 30 is followed by a mean 10-bar return higher than the unconditional mean by more than the cost assumption, in at least three of four independent sub-periods.” |
| “This system works, but only when the market is healthy.” | The regime is identified after the fact. | “…with entries taken only on bars where a broad index closed above its own 200-day average on the previous bar.” |
| “Volume confirms price.” | No claim at all; two nouns and a verb. | “…entries filtered to bars whose volume exceeded twice the 50-bar average volume produce a higher mean 20-bar return than the same entries unfiltered.” |
| “I get better fills in the first hour.” | Population of one, no measurement, no comparison. | “Across my last 200 recorded trades, the mean slippage against the arrival price on orders placed before 10:30 differs from the mean on orders placed later.” |
Notice what the rewrites share. Each one names the instruments, states a condition evaluable on a specific bar, names a horizon, names a comparison, and includes a number the result has to beat. Each is also noticeably duller than the original. That is not a coincidence: the excitement in the original sentences was mostly ambiguity, and ambiguity is what you are removing.
What this does not settle
Section titled “What this does not settle”Two limits are worth stating now, because they recur through the rest of the course.
First, a hypothesis constrains what you test, not what is true. Writing it down before you look at the data removes one bias — the one where the claim quietly reshapes itself around the result — and leaves every other bias untouched. Survivorship in your database, look-ahead in your formula and cost assumptions that are too kind will all still be there. Part 30 is devoted to those.
Second, evidence about a claim is not evidence about a system. Even a hypothesis that survives a careful test tells you about an average across many bars and many symbols. What you would actually experience is a sequence of trades with a particular order, a particular drawdown and a particular capital constraint. Parts 28 and 29 build that bridge; Part 34 deals with whether you could live through it.
An observation reports where your attention went. A hypothesis is a claim with a population, a condition, an outcome, a horizon, a comparison and a stated falsifier — the last two being the ones that are usually missing and always load-bearing. Falsifiability in markets is practical rather than philosophical: can you say in advance what would make you stop? The three ways trading claims evade that question are conditioning after the fact, discretionary overlays, and searching until a number pleases you.
A mechanism is not required to run a test, but it lowers the chance that you have found noise and it produces side predictions that are hard to satisfy by accident. Writing the claim down, dated, with a cost assumption and a reserved period, is a four-minute habit that changes what you are able to conclude later.
The next lesson takes a written hypothesis and turns it into rules — the stage where “above”, “strong” and “recently” have to become arithmetic, and where the ambiguity test decides whether you have finished.
Check your understanding
Sources for this lesson
1 verified · checked 2026-08-31
- 01AmiBroker User's Guide — Back-testing your trading ideas§ Writing your trading rulesamibroker.com/guide/h_backtest.html2026-08-31
Every technical claim on this page was checked against the official AmiBroker documentation on the date shown. Where the course disagrees with folklore, the source is how you can tell which one to trust.