Almost everyone who sets out to track their trading psychology starts by writing about how the trade felt, and almost everyone stops within a month. The writing is not the problem. The problem is that prose does not group. Sixty paragraphs about how you felt on sixty trades cannot answer the one question that would justify the effort — whether the trades you took in a particular state actually came out differently from the rest.
This guide covers what is worth recording, when to record it so the record stays honest, how to read it back, and the specific ways psychology tracking produces confident nonsense. It is descriptive throughout. Nothing here is a recommendation to trade in any particular way, and none of the numbers below are claims about traders in general — they are the arithmetic of your own log.
What tracking psychology actually means
Tracking your trading psychology means turning your own state and process into structured fields attached to individual trades, so that later you can split your results by those fields and compare. That is the whole mechanism. A state you recorded is a column; a column can be grouped by; a grouping produces two numbers you can put beside each other. A state you described in a paragraph is none of those things.
This is a narrower activity than it sounds, and the narrowness is what makes it work. You are not diagnosing yourself, and you are not trying to reach an insight at the moment of writing. You are leaving behind the minimum structured evidence that a version of you in three months can use to answer a specific question: when I logged this, what happened?
It follows that the fields have to be decided in advance and kept stable. A field you invent halfway through the sample splits your history into two incomparable halves. It also follows that fewer fields, filled in every time, beat more fields filled in when you remember — a column that is populated on 40% of your trades will mostly measure which trades you felt like annotating.
Why the feelings diary fails
The free-text psychology journal fails for four reasons, and they are worth naming separately because each has a different fix.
It does not aggregate. There is no operation you can perform on sixty paragraphs that yields a comparison. The fix is not to write less; it is to add one or two coded fields alongside the prose so the prose becomes searchable evidence rather than the primary record.
It is written after the outcome is known. A note written at the end of a losing day describes a trader who has just lost money, not the trader who took the trade. The fix is timing: the state field goes in with the plan, before the position resolves.
It is kept selectively. People write more on bad days. A record with denser entries on losing days will show that losing days had more going on psychologically, which is a fact about the writing habit, not about trading. The fix is that the field is mandatory or it is not a field.
It is read as therapy rather than as data. Re-reading your own account of a bad session mostly reproduces the feeling. The fix is that the review step asks a question with a numeric answer, which is the subject of a later section.
The state score: one number, logged before you know
The workhorse field is a single integer for your own state at the time of the trade. A 0-10 scale is the one this site's product uses, and the exact range matters much less than two properties: it is coarse, and it is logged before the outcome exists.
Coarse is a feature. Nobody can reliably distinguish a 6 from a 7 in their own head, and a scale fine enough to feel precise mostly records how recently you thought about the scale. What a coarse score reliably separates is the bottom of the range from the top — the sessions where you were tired, rushed, angry, or trading because you were bored, against the ones where you were not. That separation is the one that turns out to be worth grouping by.
Logging it before the outcome is not a nicety; it is the difference between a measurement and a rationalisation. If you fill the field in after the trade closes, the number will drift toward the result — losers get a low state score because you are now looking for the reason they lost. This is not a character flaw, it is ordinary hindsight, and the only defence against it is writing the number down while the trade is still open.
One more discipline that costs nothing: decide what the endpoints mean once, in writing, and keep the definition where you can see it. A scale whose meaning drifts across the sample is a scale that measures the drift.
The five process questions
The second layer is process, and it is better recorded as a short fixed set of yes/no questions than as an assessment. The five that cover most of the ground: did you follow the plan, did you size it the way you intended, did you honour the stop, did you exit where you said you would, and were you emotionally clean while you were in it. A third option — not applicable — matters more than it looks, because forcing a yes or a no onto a trade the question does not fit is how noise enters a small sample.
Yes/no beats a written assessment for the same reason the state score does: it groups. Five binary fields across two hundred trades give you ten comparisons you can actually run, and each one is the same shape — the trades where the answer was yes against the trades where it was no. A paragraph gives you a paragraph.
The questions are also worth keeping deliberately unflattering. The useful version of 'did you honour the stop' is answered no when you widened it once by a small amount and it worked out. A process log that records the intention rather than the action measures nothing, and the trades where a rule was bent and the outcome was fine are exactly the ones that decide whether bending it is expensive.
Grading the decision, not the outcome
A letter grade per trade — A through F, picked by you, not computed — is the field that separates decision quality from result. Its entire value depends on grading the decision as it looked at the time. A trade that followed the plan, was sized correctly, and lost money is an A. A trade that was four times normal size on an unplanned impulse and made money is an F that happened to pay.
Nobody grades that way at first, and the grade drifts toward the P&L for months. That drift is itself measurable, which is the useful part. Take fifty or more graded trades, split them by grade, and compare the results of your A-grade trades against your C-grade trades. If your self-ratings carry information about decision quality, the groups should differ. Many traders find their grades and their outcomes are barely related — that is not a reason to stop grading, it is the first honest reading you have about how well you can judge your own decisions.
Grades and the process questions overlap on purpose. The grade is your summary judgement; the five questions are the components. When the two disagree — a run of B grades where 'followed plan' is mostly no — the disagreement is the finding.
The free-text field, and the one question worth asking of it
Prose still earns a place, provided it is the last field rather than the only one. The version that survives is short and answers one fixed question: what would I need to have seen beforehand to have done this differently? That prompt produces something specific and checkable. 'Frustrated with the morning' produces a mood label you already captured in a number.
Free text is also the place where a pattern gets named before it can be counted. If the same sentence appears in five separate entries, that is a candidate for a new coded field — and adding it deliberately, from evidence, is the one good reason to change your schema mid-sample.
What you actually do with it: grouping
The review is arithmetic, not introspection. You take the state score, cut the trades into a few buckets, and put the same performance measure beside each bucket. Four buckets is a reasonable default — a bottom band, a middle, a good band, and a top band — and the measures worth reading are average R-multiple, win rate, and expectancy. Raw P&L per bucket is the least informative of the four, because it mixes in position size.
The comparison that matters is between a bucket and your own overall baseline, not against any external benchmark. There is no published number for what a trader's win rate should be in a good state, and if there were it would tell you nothing about your log. Your baseline is the only honest comparison, and the question is always the same: is this slice meaningfully worse or better than the rest of my own history?
The same grouping runs on the process fields, and the process version is usually the sharper one. Trades where the stop was honoured against trades where it was not, over a large enough sample, is a comparison with a specific and actionable answer — and unlike the state score, neither side of it depends on how well you can rate your own mood.
Sample size, and the pattern that is not there
This is where psychology tracking most often goes wrong, and it goes wrong in the direction of finding too much. Slice a hundred trades by state, by day of week, by setup, by time of day, and by whether you were following the plan, and some slice will look terrible by chance. Ten slices of ten trades will produce an apparent effect whether or not one exists.
The defensible floors are boring and worth writing down: enough closed trades overall before you look at anything at all, and enough closed trades inside a slice before that slice gets an opinion. The product this site publishes uses twenty closed trades overall and five inside any slice before it will describe that slice, and it only remarks on a slice whose win rate is at least fifteen percentage points below the baseline, or whose expectancy is negative. Those are chosen thresholds, not significance tests — there is no p-value behind them, and a fifteen-point gap on six trades is still six trades.
Two habits keep this honest. Decide which comparisons you care about before you look, so you are not selecting the worst of twenty after the fact. And when a slice looks dramatic, check what it is made of — a bucket that contains most of your trades will simply restate your own baseline back to you, and a bucket that contains one bad week is a bad week.
Sequence effects: the trade after the loss
The single most productive psychology grouping needs no self-report at all: order your trades by when you decided to take them, and compare the ones taken immediately after a loss against your baseline. Same for the ones taken immediately after a win. Both of those fields are computed from data you already have, which makes them immune to the honesty problem that limits every self-rated field.
The after-a-loss slice is the operational form of tilt, and it is where oversizing and unplanned entries usually show up first. The after-a-win slice is the operational form of overconfidence. Neither is guaranteed to appear in your log; some traders show nothing on either, and that is a real result rather than a failed measurement.
One technical detail decides whether this comparison means anything: the ordering has to be by the moment of the decision, not by when the position closed. A trader who scales into one idea over an hour will have several legs that closed after an earlier leg lost, and filing those as revenge trades measures the scaling habit rather than the reaction. If your journal orders by close time, this grouping will quietly mislead you.
Self-report is the weak link, and it is manageable
Every field in the first half of this guide is something you told yourself about yourself, which places a hard ceiling on how much weight any of it can carry. Three specific failure modes, and what actually helps:
Hindsight contamination. Covered above; the fix is writing the number before the outcome is known, and never backfilling a missed entry from memory. A blank is more useful than a reconstruction, because a blank is visibly missing while a reconstruction looks like data.
Scale drift. Your 7 in March and your 7 in September may not be the same state. Written endpoint definitions help. So does reading the score only as bands rather than as a continuous number.
The observer effect, which is not entirely bad. Rating your own state before every trade changes your behaviour — that is a confound for the measurement and arguably the point of the exercise. It does mean the earliest part of your sample is not comparable to the rest, and it means the field cannot tell you what would have happened if you had not been logging.
The weekly layer
Per-trade fields answer questions about trades. A weekly entry answers questions about weeks, and some of what you want to know only exists at that scale — whether the weeks you rated as disciplined were the profitable ones, whether volume goes up when your mental state goes down.
A workable weekly record is four scores and two sentences: discipline, setup quality, execution, and mental state, each rated once for the week, plus what you would do differently. Read back, it forms a small matrix — each of those four ratings against each of your weekly outcome measures — and the honest constraint is that it takes months to fill. Eight paired weeks is roughly the minimum before any cell in that matrix is worth a second look, and eight weeks is two months of consistent logging.
The weekly layer is also the natural home for the review itself. A per-trade review loop that runs after every trade turns into a chore and stops; a fixed weekly slot that reads the groupings from the last month is a habit that survives, and no grouping worth reading changes meaningfully within a week anyway.
What the correlations are not
A relationship between your state score and your results is not evidence that your state caused those results, and there are at least three ordinary explanations that fit the same data.
Reverse causation. A calm, orderly market produces both good setups and a calm trader. The state score is downstream of the same conditions as the P&L, and the correlation appears without your state having done anything.
A third variable. Time of day, position size, and the instrument you trade all move together with mood for reasons that have nothing to do with psychology. Before concluding that your low-state trades lose money, check whether your low-state trades are also your largest ones.
Selection. If you skip trades when you feel bad — which is the behaviour tracking is supposed to encourage — then the low-state trades that made it into the log are the ones you took anyway, which is a different population from the ones you would have taken. The measurement changes as soon as it works, and there is no clean way around that. It is a reason to hold conclusions loosely, not a reason to skip the measurement.
How TradeFlow Quantum handles it
Stated plainly rather than as a pitch, since this page is published by a journal product: TradeFlow Quantum records a per-trade state score as an integer from 0 to 10, the five yes/no/not-applicable process questions named above, a self-picked A-F grade the app never computes for you, and a free-text lessons field. The state score groups four ways for comparison — poor at 4 and below, neutral at 5-6, good at 7-8, elite at 9-10 — and a separate view tabulates score against P&L, R-multiple and win rate per bucket, per individual score, and as a scatter of every trade.
The pattern view applies the floors described above: twenty closed trades before it says anything at all, five closed trades inside a slice before that slice is described, and a fifteen-point win-rate gap or a negative expectancy before a slice is surfaced. It includes the two sequence slices — after a loss and after a win — ordered by decision time for the reason given in that section. The language is deliberately descriptive: it reports that a slice of your own history sits below your own baseline, and it does not tell you to skip those trades. Weekly review scores feed a four-by-four correlation matrix that stays unlabelled until there are at least eight paired weeks behind it.
What it does not have is worth being equally specific about, because guides like this one routinely imply otherwise: there is no sleep field, no caffeine field, and no wearable or biometric integration. Trades import automatically from 25+ brokers so that the annotation is the only manual work, but the psychology fields are all self-reported, with the sole exception of the sequence groupings, which are computed. It is $15/month or $150/year with a 7-day free trial that takes a card and charges nothing until day 7.
What tracking your psychology will not do
It will not identify the cause of anything. Every method above compares groups within your own log, and a group difference is a description. The research on why traders behave this way exists and is worth reading — the loss-aversion asymmetry and the stress work cited in the tilt entry are the two most cited findings — but neither of those tells you which mechanism is operating in your account this month.
It will not fix a strategy problem. If the underlying approach has negative expectancy, an improved mental state produces the same negative expectancy executed more consistently. Psychology tracking measures the gap between your process and your intent; it says nothing about whether the intent was any good, and mistaking one for the other is the most expensive error available here.
And it will not make the state change by itself. The record makes the pattern visible and puts a number on it. Whether seeing the number changes what you do on a Thursday afternoon when you are down on the week is not something any journal has ever been able to answer.