Dados B3 › Transparency
Transparency
Transparency — how you check the numbers yourself
Data version 2026-09-27-reapresentacao-agreg · page generated 2026-09-27 03:57 UTC. If these do not match /saude, you are reading a cached copy.
What is checked, before every publication — 466 automated tests, collected live
80 are DATA invariants (does the number tie out?) and 120 are product tests — access gate, navigation, language, AI connector. Calling all 466 'invariants' would be inflating it: the split is checkable in the list itself. Not a hand-picked highlight reel: every test below was found by introspecting the actual test files just now, the same way pytest --collect-only would. Grouped by what it protects against, not by which file it lives in.
Level 1 — Integrity: does the data tie out? (66)
I-B34— `situacao` may only be one of the four declared values — an open domain would become free text, and free text in a field that exists to separate 'confirmed' from 'unknown' reintroduces the ambiguity it was built to remove.I-B35— A company marked `com_proventos` really has dividends, and one marked `sem_proventos` has none — the stamp is recomputable from the table it describes, so it cannot lie in the dangerous direction: claiming confirmed absence where the fetch failed.I-B57— The ÚLTIMO ordering is the ORIGINAL filing and PENÚLTIMO is the restatement, and the proof is structural rather than remembered. The oldest year in the base exists only as PENÚLTIMO, because it reaches us through the following year's file, the first one we download; and the most recent year exists only as ÚLTIMO, because next year's file has not been published. If that asymmetry ever disappears, the meaning of the two columns has changed and everything depending on it must be revisited.I-B58— On the restatements page, the first column serves the ORIGINAL filing and the second serves the restatement. This was published inverted: the largest-adjustments table showed one holding company's 2012 total assets rising from 39 to 364 billion when the movement was the opposite. Counts and the absolute difference never depended on the ordering; direction did, and a wrong direction on a page whose whole argument is rigour is worse than no page.I-B01— Total assets == total liabilities (account 2 already includes equity).I-B02— Gross profit == revenue + cost (cost arrives negative).I-B02b— Pre-tax income + tax == income from continuing operations (3.09).I-B03— NOPAT = EBIT × (1 − tax rate) with the rate in [0, 45%], and ROIC is only published when average capital is positive.I-B04— The database universe is EXACTLY what the published rule produces right now — recomputed by a route independent of what the ingest wrote.I-B04b— The 20 hand-resolved companies keep their ticker in the database — enrichment must not be lost in a bulk re-ingestion.I-B27— ROA = net income / average assets, and total assets are always positive (I-B05 already guarantees 0 <= assets < 1e14).I-B06— A hole in the MIDDLE of a company's series is a loading defect — data that should be there went missing.I-B09— Gross margin recomputed straight from the RAW filing (bypassing the standardised fact) matches the published indicator.I-B26— The Piotroski F-Score is a sum of 9 binaries: it can only be an integer from 0 to 9. Anything else means a criterion leaked a non-boolean.I-B28— Current ratio recomputed straight from the RAW filing (account 1.01 of the balance sheet over 2.01, bypassing the fact) matches the published indicator.I-B30— Every company with a fact has a macro sector, and the sector is one of the 10. A new registry sector that matches no rule leaves the company sector-less and fails this test — instead of landing silently in the wrong bucket.I-B21— A dividend is a cash inflow: value > 0, always.I-B22— Type within a closed domain and ex-date in ISO YYYY-MM-DD — what the endpoints and the dividend-yield calculation assume.I-B19— The stored trailing-twelve-month profit matches the TTM recomputed HERE, straight from the annual and quarterly fact tables, bypassing the pricing code.I-B20— Every year from 2010 to the current one has prices in the database.I-B29— FCF yield = (CFO + capex) / market cap, recomputed straight from the fact (CFO 6.01 + capex 6.02), bypassing the pricing code.I-B13— The same invariant as I-B01 (assets = liabilities), at quarterly granularity.I-B10— Numbers checked by hand, one by one, against the computed output.I-B55— No company the regulator lists as non-active sits inside our universe. The page stating what this product does NOT fix publishes that intersection as zero, so if anyone ever loosens the ACTIVE-only filter on the company register, the published claim silently becomes a lie while nothing else breaks. This invariant is what turns a written promise into a checked one.I-B56— The measured size of our survivorship bias cannot quietly become zero. If ingestion of the company register fails, the page would publish 'no companies left the market since 2010', and a limitation that disappears is worse than a limitation admitted: the reader would conclude the base has no survivorship problem at all. The floor is deliberately far below the real count so only an ingestion failure trips it.I-B36— The published trail points at an account that EXISTS in that company's filing. Born from the public challenge, the first time it ran: provenance declared '3.11 (consolidated profit)' for every company, and 3.11 does NOT exist in Itau's statements — banks use a shifted chart and their profit sits in 3.09. Worse, the code varies between banks (Bradesco 3.11, Itau 3.09, Banco da Amazonia 3.13), so ingestion resolves those items by description and records the real account. The data was always right; the TRAIL was wrong — and the trail is what this product sells.I-B59— The most recent trading session in the base has to be recent. The battery had 266 checks on the CONTENT of the numbers and none on their AGE — nothing asked when the data was from. It cost 25 days in production: prices stopped on 10 August and the site kept serving them into September while the weekly rebuild reported ok every Sunday. Stale prices are not WRONG prices, which is what makes them slip through content checks; each row is still right for its own day, and it is the set that lies when it is presented as the series up to today. So this test reads no value at all, only the clock. The threshold is deliberately generous, because a failed build freezes everything, including what was fine.I-B60— The health endpoint publishes how far the series reaches, not just how many prices it holds. A count measures volume, not freshness, and it was the count alone that let the series freeze quietly: 869,232 prices looks healthy even when it has not moved in a month. This keeps the date published rather than internal, and checks that the limit shown is the limit actually enforced.I-B61— A balance-sheet account that drops about a thousandfold and comes back is flagged, and so is the return-on-equity built on it. The older detector finds the year where DOZENS of accounts are off scale, and everything downstream leans on one premise: if all of it is off by the same factor, the factor cancels in a ratio, so ROE and margins stay valid and get published. When a single account slips out alone that premise breaks — the profit is right and the denominator is not. One retailer filed equity of R$ 467 thousand between R$ 517 million and R$ 284 million, two accounts out of twenty-eight, and its ROE went out at -201.2% with no caveat when the truth was about -106%; because ROE uses average equity, the following year was contaminated too. A wrong number served as clean is the worst defect here, because the reader has no way to suspect it. An earlier version of this test only read the finished database and passed even with the flagging code deleted, so it now drives the rule itself over a synthetic series.I-B62— Equity on its way through zero is not mistaken for a scale error. A telecom holding shows equity of R$ 1.8 million between +R$ 116 million and -R$ 219 million: not a unit slip, a trajectory crossing zero on the way negative. A detector that accused that would be crying wolf, and a detector people learn to ignore costs more than none. Honest note: the sign guard was credited with protecting that case, and attacking it showed the case is actually excluded by the hundredfold minimum drop — across the whole base the sign guard currently prevents zero cases. It stays because the rule is right and data changes, but it is exercised against a synthetic trajectory so that removing it still fails this test.I-B63— The quarterly series publishes only what the quarterly filing publishes: Q1, Q2 and Q3. The regulator's interim filing never carries a standalone fourth quarter — it comes out of full year minus the nine months, an identity that closes by construction and that a direct competitor computes. We do not, because a figure we derived would enter the same list as the figures the company reported, carrying the error of two filings and erasing the line between what was filed and what we calculated, which is the line this whole database exists to hold. The rest of the test is the old rule: a ratio is only valid between accounts of the same vintage and the same consolidation perimeter, and a financial institution gets no gross or operating margin because that chart of accounts has none.I-B63b— Quarterly return on equity is twelve months of profit over the AVERAGE controllers' equity of the same twelve months, and it says so when there is no pair to average. A trailing profit against an end-of-period snapshot compares a year of flow with one photograph; the annual series already solved this with an average and declares the first year of the series. The window that matters is the SAME quarter of the previous year — not the immediately preceding quarter, which would average three months against twelve months of profit.I-B64— The decision to rebuild the database is read from disk, not held in the process's memory: a week already rebuilt successfully is not rebuilt again, and outside the early-morning window only an unfinished build from that same week is resumed. Every Sunday deploy used to rebuild the whole database, because the memory of “this week is done” was a local variable that resets with the process. Four deploys on one Sunday started four rebuilds, each killing the last, and on a single-CPU instance the visitor paid for it: a median of 1.9 seconds, a 90th percentile of 7.6 and peaks of 13 — including on the robots file, which is a constant and touches nothing. Nothing flagged it, because the panel read “running”, exactly what a healthy rebuild reads. The window itself was already promised in the loop's own docstring, which said Sunday small hours while the code fired at any hour of Sunday.I-B64b— Reading the build state never breaks the loop that it watches. A database from before this change simply has no record of the last successful build, and the behaviour then has to fall back to rebuilding — the safe side — rather than raising an error that would kill the maintenance thread. The same rule already applied to writing the state: the alarm must not be what burns the house down.I-B65— An indicator that a whole class of company never has is declared as not applicable, rather than rendered as a dash that reads as “we could not get it”. A bank's page showed sixteen years of blank return-on-invested-capital: the exclusion list named gross margin, operating margin, net-debt/EBITDA and current ratio, and forgot that one. The JSON route already said the right sentence; the page did not, and the page is what almost everyone reads. The test measures instead of trusting the list — an indicator with zero rows across all twenty-one financial companies must be declared, so the list cannot age in silence when a new indicator arrives.I-B66— The nightly price update looks for exactly the business days missing after the base's last session, up to today — and today only after the hour at which the exchange publishes its daily file; outside the night window it does not try. Prices used to enter only in the Sunday rebuild, from the exchange's ANNUAL file, which itself arrives days late: the base built on 9 September had 3 September as its last session, so the price was anywhere from zero to six days old and nothing said which. The exchange also publishes a daily file around 9 pm Brasília time; that is what now enters every night, and the price served by the API, the connector and the dump becomes yesterday's or today's close. The rule is pure on purpose, so it can be attacked without network: asking for today before 9 pm is a guaranteed 404, so is Saturday, and skipping the Monday after a Friday holiday is a hole forever.I-B66b— The nightly run ingests the daily file into a COPY of the base, checks it, and only then swaps the live base. A malformed file (minimum above close) is refused and the live base does not change by a byte; a night without a file (holiday, or not yet published) does not touch it. Never replacing good data has been the weekly build's rule from the start, and a routine that runs every night has more chances to go wrong than one that runs every Sunday. The ingestion functions commit midway, so writing into the live base would not be atomic; the copy restores atomicity, and the check is the daily gate: minimum ≤ close ≤ maximum on every new row (a shifted slice of the fixed-width layout produces a plausible number, not an error), last session advanced, nothing shrank, version stamp intact.I-B66c— The cache key of the screener, the today page and the ticker page changes when the base FILE changes, not only when the pipeline version changes. It used to be the version alone, and that does not change on Sunday: the weekly rebuild swaps the file and keeps the version, so the screener and the today page kept serving the previous week's list until the next deploy. With prices entering every night that would become a stale number every day, served from cache and looking fresh. The company page already keyed on (version, mtime) and did not suffer; the rule now lives in one place.I-B66d— Swapping the production base requires the write-ahead log of the target to be consolidated and empty; with pending frames beside it, the swap is refused. A plain file rename left the OLD log next to the NEW base. The base runs in WAL mode: if the log still holds unconsolidated frames (one open reader when the last writer closed is enough), SQLite applies them onto the new base at the next open — an old page on top of a new table. That is silent corruption no content invariant would see, and the routine that records build state in the live base is exactly such a writer. The swap now checkpoints and truncates the log first and refuses if it could not, for the nightly run and for Sunday alike.I-B67— The company profile (description, listing segment, headquarters, share registrar, auditor, control) comes from each company's MOST RECENT registration form; the listing segment is the one of the PRINCIPAL ticker; the address is the head office; the registrar is the one currently acting. An older version never overwrites a newer one, and the versioned derivative round-trips without losing a field. The first screen used to say the tax id and the regulator's code and nothing else. Comparing with a competitor made the gap obvious: what the company does, where it is, since when, in which segment it is listed, who keeps the share register, who audits, who controls. All of it was already in the files downloaded every week and simply not read. The description is the text the company ITSELF files, and the page says so; the rule choosing the version matters, because the form arrives with every version of the year and a reordering of rows must not swap the profile for an old one.I-B68— Besides the price, the nightly update asks the exchange for stock dividends and fund distributions, and the regulator for dividend filings (the future ex-date), on the same copy and before the yield is recomputed, and records how many came in. A source that is down does not bring the run down nor erase what already exists. The calendar was being born without a future: stock dividends only entered in the Sunday rebuild, so the newest announcement in the base was 18 days old and there was not one future ex-date for a stock, while funds had 138 scheduled payments. The build's rule applies here: an external source is best effort — if it fails, what was there stays, with the reason recorded.I-B69— Boot applies COLUMN migrations, not just CREATE TABLE IF NOT EXISTS — which never adds a column to a table that already exists. The migrations used to live only in the ingestion path, which runs on the weekly rebuild, so a new column could take a week to exist in production while the page depending on it kept saying 'not measured yet' with the data already in the repository. And reloading a versioned file cannot depend only on 'is the table empty?': that table was already full with the old columns, so the condition said no and the new portrait would never enter.I-B70— With the ETF table already filled, boot ADDS the rows the versioned file has and the database lacks, and leaves existing rows untouched. The old rule skipped the file entirely whenever the table had any row, so the history the backfill brought — the August 2025 anchor that enables the 12-month return — sat in the repository without ever reaching production, and the page kept showing zero funds. The house rule still holds: boot never overwrites what the build or the nightly run wrote. Completing is INSERT OR IGNORE, never REPLACE.I-B71— `dados_b3.leitura` is the stable internal reading interface: it only reads, takes the connection from the caller, imports no page or style module — and what it returns MATCHES what the company page shows, value by value and flag by flag. Sector medians are the same calculation; the company's own range uses only clean values; the universe distribution's 'fraction above X' equals a direct count. Without it, any new layer would be welded to the 1,133-line function that builds the page.I-B72— Rebuilding the database from scratch PRESERVES the snapshot of the companies that left the exchange (years of DFP filed, tickers, trading sessions), both in the table and in the re-exported versioned file. The build creates the table from the CVM registry (five columns) and rewrites the versioned file with what the table has; the snapshot comes from a manual survey that writes only to the file, and almost a third of the companies with a snapshot are not even in the registry of inactive companies. The first full rebuild after the survey rewrote the file with no snapshot, the build gate loaded that file and I-H11/I-B69 failed on every attempt — production stayed pinned to the previous data version for two days. The registry rules the five columns it knows; the survey rules the seven of the snapshot; a row only the survey knows enters whole.I-B73— The Selic target (Central Bank series 432) enters as continuous validity periods: a missing day breaks the period instead of stretching the previous rate over the gap; the fiscal-year average only comes out with the whole year covered; a source window that fails writes nothing; the versioned derived file is never rewritten with fewer days than it already covers; the boot loads it into an empty database; and the database covers every closed year with an indicator. A wrong average that looks right would decide the ROIC band of hundreds of companies.I-B74— The company page reads the indicator series, the multiples, the quarterly series, the restatements and the dividends through the data layer (first slice of the migration) — and what it receives EQUALS a direct database read, value and flag, for 20 companies of different kinds (bank, insurer, non-December fiscal year, withheld multiple). During the migration the HTML of those 20 pages, in Portuguese and English, came out byte-for-byte identical before and after, apart from the time stamp.I-B75— A registered restatement is the one of the aggregation the published figure uses that year (consolidated, or individual when the company does not consolidate), with the aggregation IN THE KEY — never again the individual statement overwriting the consolidated one. The real case, reproduced: Bradesco 2024, whose consolidated total assets matched across the two filings while the individual comparative came empty, and the old key recorded 'total assets from R$ 1.69 trillion to zero'. Measured on 26/09/2026: 9,473 → 8,660 rows, and 25 of the 356 traded companies move from a yellow to a green flag.I-B37— The revenue-growth denominator is the RESTATED comparative, not the figure as originally filed. Found by reconstructing the published sample against the raw CVM archives, after an external reviewer could audit only a handful of cases because their tooling refuses zip files. The published formula said revenue(year) / revenue(year-1) - 1 without saying WHICH version of the prior year. For one retailer the two readings differ six-fold. The comparative is the right choice for the same reason that governs the rest of the product: whoever opens a year's statement does so the following year, and the comparative shown there is already restated. Mixing the old denominator with the new numerator compares two different accounting vintages. Cases with no comparative fall back to the original figure and are stamped with a flag, so a reader can always tell which base produced the number.I-B38— A financial year whose declared currency scale is contradicted by the following filing does not publish absolute values. Found by reconstructing the audit sample against the raw CVM archives: one company was showing revenue growth of over a hundred thousand percent, because its own filing declared units while reporting thousands. The trap was diagnosing this per financial year: a filing at the wrong scale corrupts BOTH columns it publishes, its own year and the prior-year comparative, so the neighbouring good year looks broken too. The rule therefore requires both columns of the same filing to agree before calling it. And we diagnose without ever rewriting a value: a corrected number would stop matching the accounts it cites, and a number that does not tie to its own source is worse than an absent one because it looks auditable and is not. Ratios still publish, since they divide two accounts from the same vintage and the scale cancels.I-B39— No fallback flag may cover an entire class of companies. Born from a mistake made the same day the rule was written: revenue growth started using the restated comparative, with a stamped fallback for filings that have none, but the lookup matched the item name used by ordinary companies while financial institutions store it under a different name. Every single bank series fell into the fallback and none used the new rule, while the commit claimed banks were covered. The defect is dangerous because it looks tidy: each number carries a flag, and a flag reads as an explanation. Only the proportion gives it away. An exception that applies to everyone in a class is not an exception, it is the rule failing for that class.I-B40— Closing equity equals opening equity plus transactions with owners plus comprehensive income plus internal movements. Born from an audit that could not finish: a reviewer built this bridge for one company and stopped halfway, because dividends and other comprehensive income were exposed nowhere in the API, so they could reach a suspicion but not a conclusion. Their finding turned out to be a false alarm, but the hole that prevented them from confirming it was real, and this bridge is precisely the check that catches an equity error. Run across the whole base, 5,347 of 5,361 financial years close. The fourteen that do not are defects in the source, one company filing every closing balance as zero, and those are suppressed rather than published: an internally inconsistent set is worse than an absent one, because whoever checks it concludes the error is ours.I-B41— An account the regulator's file reports twice with different values is not published. The archive repeats the same account, same financial year, same column, with contradictory amounts and nothing to tell them apart: same statement group, same dates, same description. One company files profit as both one real and zero on the same line. Until now this resolved by accident on both sides: our ingestion kept the last row read, and the reconstruction script we wrote to AUDIT ourselves kept the first. The two disagreed about the same company, and that is the only reason the conflict surfaced. Arbitrary resolution does not announce itself as arbitrary; it only shows when two of your own parts choose differently, and most systems do not have two parts reading the same source by independent paths.I-B42— The scale warning travels on the FACT, not only on the indicator. Found by an external audit: the scale-divergence rule suppressed absolute indicators and multiples, but the underlying facts still shipped clean, so a utility published equity of 1.6 million where the following year's comparative says 1.6 billion. Anyone reading the raw facts or the bulk dump got a number that could be a thousandfold wrong with nothing marking it. Marking rather than suppressing is deliberate here: ratios come from these same facts and remain correct, because the scale factor cancels between two accounts of the same vintage, so suppressing would kill good data to remove bad. The general rule it closes: a warning confined to one layer is not a warning, because whoever consumes the layer below never sees it.I-B43— Every published indicator recomputes from the facts that produced it. Found by adversarial mutation testing: defects were injected into the data to measure how many the suite caught, and doubling a net margin without touching any fact went completely undetected. The product's central promise, that the published formula is the applied formula, had no invariant at all — it was checked by hand whenever someone thought to look. The drift needs no bad faith: recompute the facts and forget the indicators, which happened in this very project.I-B44— A flag the rules require must actually be present. Two mutation survivors: erasing a flag from one indicator and a shell-company warning from another went unnoticed. The existing checks verified a flag was CORRECT when present, never that it was PRESENT when due. Flags carry nearly every judgement in this base, so a vanished one returns the number to the world looking clean, which is the worst way to be wrong because nobody reading it suspects.I-B45— Impossible values are never published: non-positive prices or market caps, or a market cap beyond any plausible order of magnitude. A mutation flipped a price sign and nothing complained. A negative price is not a wrong number but an impossible one, and an impossible value passing means nobody watches that layer. The worst incident on record here was a market capitalisation of 1.67 trillion reais, impossible before it was wrong.I-B46— Return on equity and on assets also recompute from the facts, not just the simple ratios. Round one of the mutation audit found that no indicator was recomputed at all and produced the first recomputation invariant; the detection rate then hit 100%, which almost always means the attack set is too easy rather than the system being safe. Round three swapped one company's return on equity for another's: perfect shape, plausible value, normal range, only the owner wrong, and it passed. These two had been left out because they average two financial years, which is more work to reconstruct. More work is not a reason, it is exactly where defects hide, because whoever writes the test also picks the easy path.I-B47— The version stamped in the database is the version of the code that built it. A mutation replaced the stamp with an invented label and nothing complained. That stamp governs the publication gate and every claim that a user knows which state they audited. A wrong stamp is worse than a missing one, because whoever cites the version ends up citing one that never existed.I-B48— A recorded restatement must actually show a divergence. The existing check asked the opposite question, whether every real divergence was recorded; nobody asked whether every record corresponded to a divergence. The asymmetry is easy to miss and the effect is bad in both directions, since the restatement count is a headline number on the home page and an empty row inflates a transparency argument with nothing.I-B49— There is no alternative route into the batch of published facts. This invariant fixes a pattern, not a case: in one week the same thing happened five times, where a rule was applied on one path through the code and forgotten on another, and every one of them was caught by an invariant rather than by review. The cause was structural, since more than one path led into the batch and whoever added an item by a new path had to remember to reapply each rule by hand. A rule applied on three of four paths is not worth 75%, it is worth zero, because the defect picks precisely the forgotten path. The fix is not remembering better, it is making forgetting impossible, and the test is structural because a behavioural one would only fail after the next incident.I-B50— The quarterly price-to-earnings ratio equals market cap divided by trailing twelve-month profit. A mutation shifted it by 40% and nothing complained: the annual indicators were recomputed by an earlier invariant, the quarterly series had no equivalent. Same pattern that has now appeared six times in this project, a rule applied to one layer and missing from the neighbouring one.I-B51— Interest on own capital does not vanish from shareholder payouts. A mutation deleted every such entry and no invariant complained, which would have halved the dividend yield of every company that pays it, silently. This is already a declared case in the public challenge, since that instrument is booked as a financial expense rather than a distribution, so anyone summing only the dividend line understates real remuneration. There was a challenge case and no invariant: the challenge proves we can defend that case, the invariant stops it breaking unnoticed.I-B52— The sum of the quarters cannot contradict the financial year. The mutation that prompted this was the least interesting part: while calibrating the threshold, the worst cases came in at 740 times and turned out to be financial years already flagged for a currency-scale problem. The comparison was detecting the same defect through an independent path, and that became a second scale detector which reaches what the first cannot — the last filing of a series has no following comparative, so it was undiagnosable by construction. The flagged population went from 27 to 45.I-B53— An income statement filed without its cost line does not publish a clean gross margin. The investigation started in the wrong place: mutation testing flagged company-years where the quarters summed to two or three times the full year, and it was recorded as an inflated quarterly series of unknown cause. The opposite was true. The quarterly filings were right and the annual one was malformed, reporting zero cost with gross profit equal to revenue, and the annual revenue figure matched the estimated annual GROSS PROFIT rather than revenue. A hundred and seven company-years publish a gross margin of exactly 100%, which does not exist in an operating company, and the number ships clean, so anyone sorting the market by gross margin gets them at the top.I-B54— When a company's quarterly and annual filings disagree, the data says so. Of the twenty-one contradictions found, only eight had the missing cost line; the rest have real costs and still show the quarters summing to twice the year, most plausibly because the consolidation perimeter changed mid-year, which is legitimate accounting and nobody's defect. The mark is therefore descriptive rather than accusatory: it states that the company's two publications disagree without asserting which is right. The opposite temptation nearly won, declaring the quarterly wrong because the annual is the number we publish.
Level 2 — Semantics: does the account mean what the calculation assumes? (6)
I-B07— Every LATEST × PREVIOUS divergence above 0.5% in the key accounts must be recorded in `reapresentacao`.I-B24— Every bank and insurer that has a fact has both profit AND equity extracted in the most recent year.I-B15— The price used in a multiple is never earlier than the filing's receipt date — the central INVARIANT of the whole multiples block: without it, the P/E 'knows' a balance sheet the market had not yet seen.I-B16— Counter-proof: market cap recomputed outside the calculation code (price × the most recent share count AS OF the price date) matches what was stored.I-B18— The same central invariant as I-B15, in the quarterly block: the multiple's price is never earlier than the quarterly filing's receipt date.I-B32— Every indicator the database emits has a provenance entry in the DICTIONARY — the formula, the CVM accounts and the profit BASE that /indicadores now ships ALONGSIDE the number.
Level 3 — Economic: is the result possible in the real world? (8)
H-B05— Total assets never negative, never above R$ 100 trillion (catches a MIL currency scale applied twice).H-B08— |ROIC| > 200% without the `roic_extremo` flag is a defect until proven otherwise.H-B23— Same pattern as H-B08/H-B17: a payout above 5× the ex-date price is a historical price not adjusted for a split (common in older B3 data), not a parsing error.H-B25— A bank ROE outside [-100%, 150%] is almost always an extraction error (equity and profit swapped).H-B17— Same pattern as H-B08 (ROIC): |P/E| > 200 or |P/B| > 50 without a flag is a defect until proven otherwise.I-B31— The sector index carries the 10 macro sectors plus 'Market (B3)', each with a full series (same number of weeks), positive levels and base 100 at the start.I-B33— No week of any sector moves more than a stock market can move in a week.H-B14— Q1+Q2+Q3 standalone revenue <= annual revenue (2% tolerance) for the VAST majority of companies — the residual (implied Q4) must be non-negative.
Other product tests (0)
Access gate, navigation and the MCP connector — they test whether the product works, not whether a number is right, so they don't fit the 3 levels above by design. Still counted, still listed:
Reconstruct it yourself — net margin across 5 companies, 4 sectors
We don't ask you to trust the numbers. Here's the net margin of five companies in four different sectors, each rebuilt straight from its annual report (DFP) at the CVM — net income ÷ revenue, matching the published figure exactly. For ANY indicator, of any company, the whole chain down to the line in CVM's file is at /linhagem:
| Company · sector | Net income (CVM acct) | Revenue (CVM acct) | Margin =÷ |
|---|---|---|---|
| WEGE3 · industrial | R$ 6.78 bn DRE:3.11 | R$ 40.80 bn DRE:3.01 | 16.6% |
| VALE3 · mining | R$ 11.81 bn DRE:3.11 | R$ 213.59 bn DRE:3.01 | 5.5% |
| PETR4 · oil & gas | R$ 110.61 bn DRE:3.11 | R$ 497.55 bn DRE:3.01 | 22.2% |
| ITUB4 · bank | R$ 45.85 bn DRE:3.09 | R$ 387.12 bn DRE:3.01 | 11.8% |
| BBAS3 · bank | R$ 16.78 bn DRE:3.11 | R$ 319.46 bn DRE:3.01 | 5.3% |
Download any of these DFPs from the CVM, take the income and revenue accounts, divide — you get the same number. Banks use interest income as revenue (what makes sense for a bank), so the reconciliation is sector-aware; the others use sales revenue. Figures are the latest fiscal year.
Why we may differ from another site (and it's not an error)
A difference between two sites usually isn't one being wrong — it's a method choice. We disclose ours:
- Average vs ending capital: ROE and ROIC use average equity/capital (this year + last ÷ 2), not the ending balance — so they're not a naive single-year division.
- Controlling vs consolidated: we state which one each item uses.
- TTM vs annual: quarterly multiples use trailing-twelve-months earnings.
- IFRS 16, goodwill, cash, exceptional tax: each handled explicitly and flagged when it distorts.
Do the tests bite? Yes — real cases the suite has caught
A test that never fails could mean perfect data — or a weak test. These were born from real errors that slipped through, and now fail — only shown for tests that actually exist in the collected suite above:
- I-B16 · banks' P/E = 0: the CVM reports share count sometimes in units, sometimes in thousands (varying by company and year); market cap came out 1000× too small and P/E was zero. I-B16 (price × shares recomputed) caught it — we fixed 661 annual and 2,027 quarterly multiples.
- I-B24 · the vanishing profit: the profit/equity account varies by bank (Itaú 3.09, BB 3.11); if the label search fails, the number vanishes silently and ROE is born wrong. I-B24 makes that hole fail.
- H-B23 · the 800% dividend yield: old B3 dividends carry a price not adjusted for splits; without H-B23 the yield would look absurdly real. The flag keeps the record and warns.
- I-B20 · the missing COTAHIST 2023: on the first deploy one year of prices failed to download silently; the API shipped with ~50k fewer prices and the suite passed, because nothing checked coverage. We added I-B20 — a year without prices now fails the build loudly.
- I-B33 · Oil & Gas at −89.6% in one week: COTAHIST isn't split-adjusted; a 10:1 split read as a −90% weekly return and stayed in the cumulative return forever, with nothing failing (a blind amplitude filter didn't fix it either — it just flipped which direction was wrong). I-B33 (no sector week beyond ±35%, the ruler is the exchange itself) stops the build.
- I-B31 · 364 of 865 sector-index weeks frozen: extending the series back to 2010, a sector with no company carrying a market cap that week turned into a factor of 1.0 — a flat line that the cumulative return read as 'market stood still', and the 'since 2010' return came out fictional (+527%). I-B31 fails on any week repeated to the cent.
- I-E05 · EMBRAER's 2010 P/E priced with 2025 data: the point-in-time price lookup fell forward with no ceiling whenever the current ticker had no price at balance-sheet time (ticker change, share-class migration); the multiple came out priced 15 years into the future, in 22% of annual and quarterly rows, silently. The 45-day window (and I-E05) blocks it: no session inside the deadline, no multiple computed.
- I-E04 · AZUL at R$ 1.67 quadrillion market cap: the balance sheet carried 54.7 trillion shares from the judicial recovery issuance, but the price used was already on the post-reverse-split (1:150,000) basis — two ends measured on different bases, market cap wrong by orders of magnitude. I-E04 (market cap never above Brazil's GDP) fails the build on that absurdity.
How far back each block of data goes
No asterisks here: this is coverage year by year, counted right now. Indicators (ROE, ROIC, margins, growth) come from the filings alone and cover the whole series. Multiples (P/E, P/BV, EV/EBITDA) need a price and a share count — and the share count comes from CVM's reference form, which the further back you go the fewer companies filed in a usable format. The gap between the two columns is a source limit, not an ingestion hole. If you look for an old P/E and do not find it, it is because nobody has it — not because we hid it. The same numbers as JSON: /cobertura.
| Year | Companies with indicators | Companies with multiples |
|---|---|---|
| 2025 | 438 | 290 |
| 2024 | 444 | 296 |
| 2023 | 442 | 298 |
| 2022 | 429 | 295 |
| 2021 | 423 | 296 |
| 2020 | 410 | 274 |
| 2019 | 371 | 164 |
| 2018 | 326 | 164 |
| 2017 | 317 | 162 |
| 2016 | 311 | 152 |
| 2015 | 303 | 143 |
| 2014 | 301 | 139 |
| 2013 | 292 | 130 |
| 2012 | 291 | 111 |
| 2011 | 284 | 97 |
| 2010 | 277 | 34 |
Live coverage
456 companies · 876,076 price points · 8,654 recorded restatements · last refresh 2026-09-27 00:20:28. Full live counts at /saude.
Sources: CVM (open data, ODbL) and B3 (COTAHIST). Not affiliated with B3 or the CVM. Not investment advice.
Numbers on this page are live. Data version 2026-09-27-reapresentacao-agreg · page generated 2026-09-27 03:57 UTC. If this does not match /saude, you are reading a cached copy.