Teardowns takes a claim apart. A spec-sheet line, a keynote number, a default setting nobody questions, a marketing phrase that sounds quantitative but is not.
The method is consistent: we work out what conditions would make the claim true, reproduce those as closely as we can, then run the same test under the conditions a person actually lives in, and publish both numbers. Usually the claim turns out to be technically accurate and practically misleading, which is a more interesting result than either "true" or "false" and is much harder to find written down anywhere.
This is the pillar most likely to be mistaken for criticism, so the standard for it is the strictest on the site. Run counts, variance and ambient conditions are stated. Where our result differs from a published figure we say what we think explains the gap rather than implying bad faith. Numbers we cannot stand behind do not get published, even when they would make a better headline.
The most common finding in this pillar is not that a claim is false. It is that a claim is precisely true under conditions nobody actually experiences, and that the gap between those conditions and ordinary use is much larger than the marketing implies. A battery figure measured against a defined workload, a speed measured on an idle connection, a core count that treats two different core designs as interchangeable: none of these is dishonest, and all of them describe something other than what a reader assumes.
That is a more useful result than a verdict, and considerably harder to find written down. It also means these articles tend to end without a recommendation. The purpose is to establish what a number means and what it does not, and to hand the reader enough method to check it on their own hardware, rather than to tell them what to buy.
Every teardown states the run count behind each figure, and where a measurement was taken only once it is described as an observation rather than a measurement. Where readings varied between runs, the spread is published next to the mean instead of being averaged into a cleaner-looking single number. A figure with a wide spread is itself a finding, and hiding it would misrepresent how reliable the result is.
Nothing in this pillar is adversarial toward Apple, and nothing in it is deferential either. A claim that survives testing is reported as surviving, which happens regularly and makes for a duller headline. The interesting cases are the ones where a figure is accurate and the impression it creates is not, and separating those two things is most of the work.
Where our result disagrees with a published figure, the article states our conditions in full and offers what we think explains the gap. Sometimes that explanation is that the official number was measured against a workload nobody runs. Sometimes it is that our own test was measuring something subtly different, and when that turns out to be the case we say so rather than quietly dropping the article.