Dados B3 › What we got wrong ·
Português
What we got wrong, and fixed
These are 68 defects we found in our own database. Each one is here with the date it was found and the invariant that now keeps it from coming back — click the code to see the test on the transparency page.
“Trust us, the data is good” is worth nothing coming from whoever sells the data. This list is worth something, because a shop that hides its mistakes does not publish a table of them.
It is not curated. The entry rule is mechanical: the invariant tells a dated story. I do not pick what shows up — for a defect to leave this page its test would have to be deleted, and that fails the battery before anything is published.
2026
- 2026-09-07
I-N43
An English page is either in English, or it says why part of it is not. Sweeping all seventy-five pages for words that exist only in Portuguese turned up eight, and seven were false positives of the most instructive kind: a company's legal name and the regulator's own account labels. Proper nouns and primary-source labels are not translated, because translating them would invent data the source does not have. The eighth was real: the audit-history page translated its frame and left seventy findings in Portuguese with no explanation. Translating those would be worse than leaving them — a rewritten audit record stops being a record, and the difference between what the auditor said and what we say they said is the whole point of keeping the table. So the page declares it, which is what this codebase does with every limitation. - 2026-09-06
MCP10
The tool list published in llms.txt is generated from the connector's own catalogue, not kept as a parallel list by hand. The hand-written one was already wrong: it announced ten tools while the connector exposed fifteen, so a third of the product was invisible to any AI reading the file that exists precisely to introduce it. Two lists far apart age in silence; this one is derived, and a new tool that does not reach the front door fails the build. - 2026-09-06
I-N42
The what-we-got-wrong page is generated from the invariants' own docstrings, not from a hand-kept list, and every dated defect in the code appears on it. Each invariant here was born from a real defect and keeps the story, with its date, in its docstring: sixty-five of them, written over months, none of which had ever left the code. What this test holds is not the page but the impossibility of curating it — a hand-written list would allow choosing what to show, and the temptation would be to omit the ugliest defect exactly when it is the most instructive. The page is a function of the docstrings: if a dated invariant exists in the code, it must be on the page. - 2026-09-06
I-N41
A company page carries the quarterly series, every figure equal to the database, with no fourth quarter, in both languages. The annual series only exists once the fiscal year closes, so a company that filed its second quarter in July still showed its last full year, and all 455 pages stayed identical for twelve months. The quarterly data had always been in the database; it had never reached the person reading. A page that changes every quarter is one a search engine and an AI have reason to revisit; one that changes once a year is not. The fourth quarter stays out here for the same reason it stays out of the API: the interim filing does not publish it on its own, and deriving it would put a figure we computed in the same table as the figures the company reported. - 2026-09-06
I-N40
The methodology index answers in JSON to whoever asks, carrying the title of each page. The MCP connector publishes a methodology() tool that called this index and parsed the result as JSON — and the result was always HTML, so any AI calling that tool got a parsing error instead of an answer. It survived because reading one page always worked; only the index was dead, and nothing on our side exercised that path. The lesson is not that JSON was missing: a published contract with no invariant exercising it is a promise, and the tool sat in the list, fully described, with no one on our side ever calling it. - 2026-09-06
I-N37
The what-changed page opens under any Accept header, is in the sitemap and in the panel's explicit list, and declares both its window and how far each dataset reaches. It exists because after 820 open pages and a screener the site was an excellent dictionary with no reason to come back tomorrow; this is the recurring reason, built entirely from our own data. Each source lags differently, so the page shows the lag instead of hiding it. - 2026-09-06
I-B64
The decision to rebuild the database is read from disk, not held in the process's memory: a week already rebuilt successfully is not rebuilt again, and outside the early-morning window only an unfinished build from that same week is resumed. Every Sunday deploy used to rebuild the whole database, because the memory of “this week is done” was a local variable that resets with the process. Four deploys on one Sunday started four rebuilds, each killing the last, and on a single-CPU instance the visitor paid for it: a median of 1.9 seconds, a 90th percentile of 7.6 and peaks of 13 — including on the robots file, which is a constant and touches nothing. Nothing flagged it, because the panel read “running”, exactly what a healthy rebuild reads. The window itself was already promised in the loop's own docstring, which said Sunday small hours while the code fired at any hour of Sunday. - 2026-09-06
I-B63
The quarterly series publishes only what the quarterly filing publishes: Q1, Q2 and Q3. The regulator's interim filing never carries a standalone fourth quarter — it comes out of full year minus the nine months, an identity that closes by construction and that a direct competitor computes. We do not, because a figure we derived would enter the same list as the figures the company reported, carrying the error of two filings and erasing the line between what was filed and what we calculated, which is the line this whole database exists to hold. The rest of the test is the old rule: a ratio is only valid between accounts of the same vintage and the same consolidation perimeter, and a financial institution gets no gross or operating margin because that chart of accounts has none. - 2026-09-05
TR35
The bulk-export index states how many rows come without a traded code. Its note used to promise that every line carried the ticker and the CVM code, so the datasets would join against your own base — false for one row in six, 11,739 of 71,085 in the indicators set, across 96 companies. It is not a broken join: those are companies registered with the regulator that have no traded stock, and not one of them has a single trading session, so there is no code to carry. The data is right; the sentence promised more. The cost falls on whoever consumes it and is invisible — joining on ticker drops those rows with no error at all, and an outside audit that found the gap concluded it was a failing join, which is itself the symptom of an undeclared limit. Same defect as the price-session wording fixed a day earlier: a claim above what the data delivers, always in our favour, and the same fix — publish the size, computed from the base. - 2026-09-05
I-N36
The about page names the person who builds the site, in both languages, with the credentials that speak to this product. It used to explain the operating company and say nothing about who makes it — and a site with no visible owner is the first thing a search engine or an AI discounts. A reference has a name. - 2026-09-05
I-N34
Every company page links to the other companies in its CVM sector, and every fund page to the other funds in its segment — all of them, none of itself. The 820 pages opened in September were born as islands, reachable only from the hub or the sitemap, and a page no other page cites is the last one a crawler visits; Search Console showed forty of them detected but not indexed. The grouping is the regulator's own classification, not ours. - 2026-09-05
I-B61
A balance-sheet account that drops about a thousandfold and comes back is flagged, and so is the return-on-equity built on it. The older detector finds the year where DOZENS of accounts are off scale, and everything downstream leans on one premise: if all of it is off by the same factor, the factor cancels in a ratio, so ROE and margins stay valid and get published. When a single account slips out alone that premise breaks — the profit is right and the denominator is not. One retailer filed equity of R$ 467 thousand between R$ 517 million and R$ 284 million, two accounts out of twenty-eight, and its ROE went out at -201.2% with no caveat when the truth was about -106%; because ROE uses average equity, the following year was contaminated too. A wrong number served as clean is the worst defect here, because the reader has no way to suspect it. An earlier version of this test only read the finished database and passed even with the flagging code deleted, so it now drives the rule itself over a synthetic series. - 2026-09-04
TR34
Every public route WITH A PATH PARAMETER opens under any Accept header and is counted by the panel. TR16 audits public HTML pages but drops every parameterised route by construction and only looks at routes declaring an HTML response class, so the company pages, the fund pages and the 500 audit samples were all invisible to it. The samples served raw JSON to */* and were measured by nothing at all: the panel read zero for them since forever, and that zero meant not-measured, not nobody-came. This test also refuses to let the list of parameterised routes go stale, and proves the data-route label is honest rather than a hiding place. - 2026-09-04
I-N32
Filters stack rather than replace, and each one can be removed on its own. A screener whose second filter silently replaces the first looks like it works — it returns rows and raises no error — and does not do the job; only counting catches it. Each active filter also carries a link that drops just that one and keeps the rest. - 2026-09-04
I-N31
The screener table is rendered by the server, with no JavaScript. Building the table in the browser is the natural choice and would hand a crawler or an AI an empty page — and the 820 public pages opened in September are worth something precisely because the content is in the HTML. So sorting is a link and filtering is a GET form: the crawler walks exactly what a person sees, and every slice is an address. - 2026-09-04
I-N30
No URL in the sitemap answers with a noindex tag. The two are opposite instructions, and for months 500 audit-sample URLs carried both: they were listed in the sitemap on the sound reasoning that a tool which only opens indexed URLs stalls otherwise, while the page itself said noindex — so it could never be indexed, the tool stalled anyway, and the entries took up 36% of a sitemap on a site where Google already reported forty pages detected-but-not-indexed. Each half was defensible alone and nothing looked at them together. The fix was to drop the samples from the sitemap rather than drop the noindex: the seed hub is indexable, sits in the sitemap and links all five hundred, so the tool reaches the hub and follows a link, while five hundred generated near-duplicates stay out of the index where they would compete with the real pages. - 2026-09-04
I-N29
The page and the JSON of an audit sample carry the same cases. Serving one sample in two formats risks them drifting apart, and then a seed would have two versions — which is exactly what TR24 prevents between the two URLs, while nothing prevented it between the two formats of one URL. An earlier draft of this invariant demanded the page be served to */* by analogy with the fund pages; the battery refused it, correctly, because there */* got a 401 with no content while here it gets the whole artefact, and that JSON is the citable proof the published protocol tells auditors to request. - 2026-09-04
I-N28
Without a key, a fund page returns the PAGE under any Accept header. The fourth time in this family, and this one was self-inflicted hours after the page was written: the route first negotiated on an explicit text/html, so Googlebot got the page while */* — the default of curl, GPTBot and ClaudeBot — got a 401. Search Console refused the indexing request for that URL and the reason was exactly this. The rule is the KEY, not the Accept: whoever sends a key wants the feed, and whoever does not has no access to the JSON anyway, so a 401 only hides public content. - 2026-09-04
I-N25
A fund page answers 200 in HTML WITHOUT a key, on the SAME URL that serves the JSON. The format comes from the Accept header: a browser reads, an agent consumes. Before this the hub linked each of its funds to a gated JSON, so a person who clicked a name got a 401 — explained, but shut — and to a search engine or an AI the fund did not exist. If anyone flips the order, gating the HTML or serving JSON to browsers, those links go back to being a closed door and the funds go invisible again. - 2026-09-04
I-B59
The most recent trading session in the base has to be recent. The battery had 266 checks on the CONTENT of the numbers and none on their AGE — nothing asked when the data was from. It cost 25 days in production: prices stopped on 10 August and the site kept serving them into September while the weekly rebuild reported ok every Sunday. Stale prices are not WRONG prices, which is what makes them slip through content checks; each row is still right for its own day, and it is the set that lies when it is presented as the series up to today. So this test reads no value at all, only the clock. The threshold is deliberately generous, because a failed build freezes everything, including what was fine. - 2026-09-03
PR06
The plans page states who the charge comes from. Forty-six checkouts were opened and none was paid, and at the last step the customer met a company name the site had never mentioned. The public name at the payment provider is now the product's, but the registered entity is still another one and it is the entity that appears on the invoice — so the page says so BEFORE, instead of leaving the discovery for the moment of the card. - 2026-09-03
PR05
The free-key button leads to a form, not to the payment provider. The free tier is a zero-value subscription, and the path to it used to be the same checkout as the paid plan — no card requested, but wearing the face of a payment form. The funnel measured the cost: 47 people opened it and 16 finished. Asking someone to cross a billing screen to collect something free is friction with nothing on the other side. - 2026-09-03
MED15
Every authenticated call goes through the usage counter, and the counter exists. The strategy became 'make the people already using it dependent on the product, then monetise' — which requires knowing whether anyone comes BACK, and that was exactly the question with no instrument: the counter lived in an in-memory dictionary, wiped on every restart. We knew sixteen keys had been created and nothing about what happened next. Third time the same lesson appears: the channel we bet most on was the only one without an instrument. - 2026-09-03
I-N22
A company page answers 200 WITHOUT a key and is in the sitemap. That is the whole point of the page: the site was invisible three ways at once — the sitemap held 568 URLs of which 503 were audit samples and none was a company, a search for a well-known company's ROIC returned seven competitors and not us, and an AI asked for a bank's ROE hit a 401 and answered with another site. If anyone puts a gate here all three holes come back silently, because the page still exists and nobody can reach it. - 2026-09-03
I-N20
Any page using the fact-page frame is registered in the panel's explicit list. The panel returns the thirty most-visited pages and its last row had twenty-eight visits, so a fact page born with five simply vanishes from it — and 'did not appear' reads as 'nobody visited' to whoever is looking. The fix was listing each one explicitly, which only works while the list stays complete; whoever adds the twelfth page and forgets to register it finds out here rather than a month later, staring at a zero that was never a zero. - 2026-09-02
I-H11
A database missing the derived tables gains them at BOOT, not through a full rebuild. Covenants and the excluded-companies list depend on nothing the rebuild recomputes: they come from files versioned in the repository. Taking the official route of bumping the data version would cost around twenty-six minutes of reconstruction in a build that has already died of memory exhaustion and, on one occasion, took down an external audit in progress. The test loads into an empty database and checks it loaded, then runs again and checks it did NOT reload, because overwriting what a build placed would be the boot overruling the build's authority. - 2026-09-01
TR33
The comparison page states when it was checked and admits where we lose. It had been wrong in our own favour for three weeks: the table said a competitor had no MCP connector, only a partial methodology, and a quarterly price around forty reais. In fact it now ships a native MCP connector installable in one line, publishes a methodology that cites regulator account codes, and charges essentially our price with a free tier shaped like ours. That page is the SECOND most fetched by AI assistants, which makes it the worst possible place to be wrong in our own favour, and simply re-dating it without re-checking would have preserved the error - the word verified only means something if someone verified. The test attacks the shape a marketing document takes when it pretends to be a comparison: it requires the section listing where each competitor beats us to exist, requires a concrete admission (one of them has price history going back twenty-four years further than ours), and requires the competitor's connector to keep being acknowledged. - 2026-09-01
I-N19
Someone arriving without a key gets the PAGE, whatever their Accept header says. This was a discovery defect found in production: the first version negotiated on an explicit text/html header, so a browser got the page while the wildcard header - the default for curl and for several automated fetchers - got a 401, as did a request with no Accept header at all. The page is listed in the sitemap and in the file we publish for AI agents, so anyone arriving through either got a closed door on the most differentiated page we have, and a 401 reads as nothing-here rather than wrong-format. This is the third time in the same family: the connector endpoint answering 406 in a browser, the audit route missing from the AI index, and now this one - the pattern is always a door that only opens for someone who already knows how to knock. The correct rule is not the Accept header but the API KEY: whoever sends a key wants the feed and existing integrations must not break, while whoever sends none has no access to the JSON anyway, so answering 401 merely hides public content. - 2026-09-01
I-N16
The sustainability-reporting page states that the obligation was REPEALED, with the date. The idea arrived with an out-of-date premise, and that is what made it worth publishing: it came as companies being required to report from 2026 with the first reports in 2027, which was true until 29 May 2026, when a new resolution repealed the requirement outright rather than postponing it. Almost everything written in Portuguese on the subject predates that repeal and still says it is mandatory, so a page that corrects information an assistant would otherwise repeat is exactly the kind of page that gets cited. The page also declares what we do NOT have: we ingest annual and quarterly financial filings, not sustainability reports, so this is regulation rather than our own measurement. Without that admission the page would imply we measure environmental and governance data, which would be more useful to us and less true. The test attacks by removing either the word that carries the correction or the admission of the limit, and it also requires the sources to appear on the page, because a rules page without sources is an assertion. - 2026-09-01
I-N15
The bank chart-of-accounts page proves that an account code does not define an account. The fact it carries: one large state-owned bank reported profit under one code through 2019 and under a different one from 2020 onward, while the largest private bank stayed on the original code for the entire series. Anyone scraping the regulator with a fixed code gets the private bank right, gets the state-owned one wrong from 2020, and receives NO ERROR at all: they get an empty value, or another line's number. The stable case being the most famous bank is what makes this dangerous, because it is the one everybody tests a scraper against, and it passes. Seventeen of twenty-one institutions changed their equity account, most of them in the same year. This is not our theory: an external auditor found the mirror defect in our own output, where we declared one account and used another, and that correction became its own invariant. This page publishes the map that episode showed was missing, and the test attacks by checking the contrast survives. - 2026-08-31
I-N14
The data-quality page lists what we do NOT publish, and why. Writing it fixed a defect before anything was published: while assembling the breakdown of the eight thousand withheld indicators, over a thousand turned up with a NULL flag, meaning they were withheld with no stated reason. On a page whose whole subject is why a figure is missing, thirteen percent of 'I do not know' is the page contradicting itself. The cause was zero revenue on the income statement, which happens in holdings whose result comes from equity income and in companies with no operations that year: the withholding was correct and the reason was simply not declared. Those became explicit flags and the silent cases fell to fifty-three. Withholding without saying why is half the house rule; the other half is saying it, and that is the half a public page enforces. The test also caps the silent cases, because without a ceiling the next rule that withholds without declaring would go unnoticed exactly as these did. - 2026-08-31
I-N13
The restatements page is public while the per-company feed stays behind a key. This is the first fact page, and the rule behind it comes from measurement: over thirty days the AI assistants fetched the home page 209 times, the comparison page 15, the multiples methodology 14 — and the TWENTY conceptual guides added up to 7. An assistant already knows what a price-to-earnings ratio is; it fetches whoever answers what it does NOT know. Restatements are the one product dataset no competitor publishes, because the regulator overwrites the old version of a statement and we keep both. The split mirrors the rankings page: the AGGREGATE is public because it is acquisition, the per-company detail is paid because it is the product. Publishing the whole feed would give away what sustains the paid tier; hiding the aggregate would hide precisely what differentiates us. - 2026-08-30
PG01
No module that tests DATA is left outside the publication gate. The gate was trimmed for a real reason and trimming is dangerous: four consecutive rebuilds died without recording any failure, the whole service dropping at around fifty-seven minutes, and the navigation module accounted for 185 of the battery's 220 seconds, crawling every page of the site through a test client while the freshly built 1.8 GB database sat open beside it. A broken link does not corrupt data, and it was blocking correct data from going live. But the next temptation is obvious, and this test exists to block it: to keep cutting until the gate is fast and empty. A gate that rejects little is not a cheap gate, it is an ornament. The check also verifies that every module the gate names actually exists on disk, because pytest given a missing target fails, and that would break every rebuild. - 2026-08-30
AUD15
The excluded depreciation lines are published, not only the ones that were summed. Without this the new rule was not falsifiable from outside: we published which accounts ENTERED the depreciation figure and nothing about which stayed out, so an auditor wanting to attack the exclusion of debt-issuance cost — a change that moved 232 EBITDA values — could see neither what was excluded nor why. The rule matches the NATURE of the label rather than the account code, since the same numeric code is transaction cost at one company and amortization of a sales stand at another, and a rule like that can only be defended by showing the label that triggered it, which means it can only be ATTACKED the same way. The auditor named three companies as targets; checking them showed the rule does not even fire at two of them, so the route exists for him to find the cases where it does fire on his own, rather than depending on the audited party to point at the battlefield. - 2026-08-30
AUD14
The published population declares an arithmetic that has to reconcile against a different route. The external auditor had just proved, independently, that the sample is the top of the ordering over the population file — but the SIZE of that population was still a bare assertion: the health endpoint published one indicator count and the population file delivered a smaller number of keys, with a gap of several thousand that appeared nowhere. He could verify that the sample was the top, and could not verify that the population was the population. Publishing the count of suppressed indicators closes the arithmetic across the two routes. This does not remove the dependency on us, since both numbers come from here, but it replaces an assertion with a SUM, and a sum that has to reconcile is attackable: dropping a key from the population now requires editing the health endpoint too. The attack truncates the population file without touching the other route. - 2026-08-30
AUD13
An auditor can redo the ORDERING, not merely the individual scores. This is the finest objection we have received: having checked every published sample score and found them all correct, the auditor observed that this proves the key reproduces the scores shown, but not that those are the smallest scores in the whole population, because he never received the population to sort. He is right, and the distinction is subtle — checking each published score detects a fabricated key, but it cannot detect OMISSION. A case with a smaller score could have been left out and nothing in what he received would reveal it, so the claim that the sample is the smallest N rested on our word. A new route publishes the population keys, and this test does exactly what he would do: download them, score all of them, sort, and compare. Its first version had the same blind spot one level up — an attack that dropped one arbitrary line PASSED, because an arbitrary line is almost never in the top N — so the count of published keys is now checked against the table itself. - 2026-08-30
AUD12
For a bank, the generic list of source accounts is not presented as though it applied. The sixth external auditor found that one bank case declared one pair of account codes while the accounts actually used were a different pair entirely: declared did not match used. He classified it precisely as an audit-trail defect rather than a numerical one, and refused to turn a documentation inconsistency into a claim that the ratio was wrong — the accounts actually used are the correct banking structure. But 1,407 indicators, 2.2 percent of the base, were published that way, and the placement is the worst part: the banking chart of accounts is exactly the surface our own prompt tells auditors to attack. A field called declared that does not describe what was used looks auditable and is not. The chart varies even between banks, so swapping in another fixed list would only change whose statement is wrong; the honest output says the generic list does not apply here and points at the field that does. - 2026-08-30
AUD11
The sample cache makes a retry instant and never serves a sample built from an older database. Six external attempts have now ended on transport rather than arithmetic: measured in production, the same route ranged from under two seconds to nearly fifty across six consecutive calls, with one exceeding a sixty-second ceiling, while locally it answers in 136 milliseconds. That is not the algorithm, it is contention on a half-CPU instance serving a 1.8 GB database from network-attached disk. Every audit tool retries the URL after a timeout, and the three that tried reported the same symptom under different names: a gateway error, a cache miss, an unavailable endpoint. Caching turns the second attempt into five milliseconds. The danger of the cache is precisely the defect it imitates: one auditor spent twelve reads looking at a frozen snapshot from the previous day and reported that the correction was not live. If WE served a stale sample after a rebuild we would be the cause of that, with the aggravation that the number would be wrong rather than merely old — so the database version is part of the cache key, and this test attacks by removing it. - 2026-08-30
AUD10
The audit protocol declares the conflict of interest instead of pretending it away. The prompt used to open with a sentence saying the person asking is not the owner of the database and has no stake in the result — and the people who paste this prompt are overwhelmingly us. The fifth external auditor caught it: the audit opened with a claim of independence and the outcome arrived in the first person, ours, admitting fault and announcing the fix. His finding survives the conflict, because the account he checked sits in an audited financial statement rather than in his trust of whoever asked. But a protocol whose entire thesis is honesty cannot begin with a false sentence, and that sentence was mine, written to sound neutral. Declaring the tie is stronger than hiding it: an auditor who knows who asked calibrates his skepticism, while one who finds out afterwards discounts the whole result, and is right to. - 2026-08-29
MED12
No page hand-writes a number the database already knows. This is the gap in the stamp rule, found while checking the AI connector page: the stamp rule only binds pages that USE a placeholder, and the page that teaches an AI how to connect escaped it by using none at all, carrying a typed 400-plus companies and sixteen years when the real figures were 456 and seventeen. The damage is specific and about as bad as it gets for that particular page: an AI reading it to answer questions about the product repeated numbers SMALLER than the truth, so we were understating ourselves on the one page whose audience is precisely the reader who will not check. Sixteen occurrences across seven files, English guides included, while the template helper already said in writing that a changing number must be a placeholder and never typed into the HTML — another rule written and never enforced. What the check deliberately does NOT flag is a third party's number: the comparison table cites a competitor's coverage, and that is a fact about them, while our own row in the same table uses the placeholder. A checker that rejected the competitor's figure would only teach people to route around it. - 2026-08-29
MED07
A page that publishes a live count also says when that count is from. The rule was already written and applied to only half the pages: the stamp helper says it exists for any page carrying a live number, and two of four had it. An external reviewer then read a cached copy of the methodology page showing one test count while the live home page showed another, and concluded a number had been hardcoded in the HTML. Nothing was hardcoded, and both pages serve the same number today: the actual defect was that the methodology page could not prove how old it was, while the transparency page, which does carry a stamp, let the same reviewer spot his own cache and not report it as a data error. A rule declared and half-applied is the failure this project exists to hunt, so it is now structural: any file with a live-count placeholder is rejected until it carries the stamp. - 2026-08-29
MED04
The build alarm lights up on the fact, not on the instant. It went silent through seven consecutive failures in production because it keyed on the state being 'failed', while the retry loop rewrites the state to 'running' at the start of every new attempt: the alarm switched off for the whole duration of each build, and a build retrying hourly is 'running' most of the time. The health endpoint was answering with seven failures, a timeout error and a null alert on the same line while production served two-day-old data. An alarm that blinks is not an alarm. - 2026-08-29
AUD09
Amortization of debt issuance cost is kept out of the depreciation figure that feeds EBITDA. This is the first data finding to come from an external auditor since the earlier successful one, and it started from a single case worth 0.34 percent: one company added a line labelled amortization of transaction cost to its depreciation and amortization, and the reviewer flagged it as a methodological caveat rather than declaring a discrepancy, because he could not open the source file to show the account. He was right, and the cause was systemic: the matcher accepted any label containing the word amortization. Measured across the base: 279 company-years affected, 226 inflated EBITDA figures, 107 of them by more than one percent and 35 by more than five, with the worst at nearly 39 percent. EBITDA is operating earnings plus depreciation and amortization OF ASSETS, and financing cost is not that, so this is a correction rather than a mark: the rule about marking instead of rewriting applies when the source is ambiguous, and a line that says cost of raising debt is not ambiguous. Half of this test guards the opposite error: amortization of the fair-value step-up from a business combination IS asset amortization and belongs in EBITDA, and the first exclusion pattern I wrote would have silently removed it from 89 lines. - 2026-08-29
AUD08
The audit sample route is named in llms.txt and the home page links to the audit protocol. The fifth external attempt produced the sharpest diagnosis yet: the reviewer picked its own seed, composed the URL, and its own reading tool REFUSED it, because that tool only opens URLs that already appeared in the conversation or in a search result, and no search engine had indexed the domain, so there was no path by link. Checking that turned up two defects of ours: the home page did not link to the audit protocol at all, and llms.txt — the file we publish precisely so that AI agents can find their way around — never mentioned the audit route, which lives only in the sitemap, a file reading tools do not consult. Publishing a door and hanging no sign on it is the same as having no door, for anyone who does not already know it is there. The file now also lists a concrete indexed URL, because a tool that cannot compose a URL cannot use a template. - 2026-08-29
AUD07
A browser gets a page; an agent that asks for JSON gets JSON. This is the fourth external audit attempt stopped by the same kind of door: the reviewer read the 47 KB home page and could not read the audit sample route at all. The route answered correctly in under two seconds — the problem was that it was 142 KB of raw JSON, a format browsing tools truncate or refuse. Half that weight was my own defect from the same day: the coordinate block repeated all twenty-odd accounts of the company for an indicator that consumes three, and a coordinate for an account the formula never touches does not help anyone audit, it just pushes what matters past the tool's reading limit. Trimming brought it to 60 KB; serving HTML closes it. Same remedy as the connector endpoint that used to answer 406 in a browser: someone arriving by browser is not speaking the protocol, they are trying to READ. - 2026-08-29
AUD05
The methodology says what happens WHEN the reporting basis changes. I claimed the rule was undocumented and I was wrong: I searched the HTML shell and a content directory that does not exist, while the text lives in the sources-and-standardization page and has been live all along. Concluding absence from my own bad search is the same mistake an external reviewer made with a cached page on the same day. What was genuinely missing is the CONSEQUENCE at the boundary: that comparing two years there compares two different consolidation perimeters. This test now requires the rule, the flag name, and the concrete case to all appear in both languages, because an abstract rule with no example is exactly what let this boundary go unnoticed for months. - 2026-08-29
AUD04
End-to-end check on the route external auditors actually use, which builds its coordinates by a different code path than the lineage does, so one passing does not imply the other. It found a real defect on its first run: the statement of changes in equity carries an extra dimension, so the three-column filter isolates one row in the balance sheet and income statement but returned SIX there, one per equity column. We would have published an ambiguous coordinate for every item sourced from it, the auditor would have picked one of six in the dark, found a number that genuinely exists, and reported a discrepancy against a correct database. - 2026-08-29
AUD02
The instructions do not promise a check that the source makes impossible. The first draft said to compare the archive checksum and that it must match. Running our own instruction before publishing showed it failing: the regulator REPUBLISHES these archives, and the one for 2024 was modified two days after we read it. Any auditor downloading today would get a different checksum and conclude we had tampered with the data. Publishing a check that fails by design is worse than publishing nothing, because it carries the appearance of rigor while manufacturing a false accusation against our own database. The checksum is still published, but as a statement of which snapshot we read, never as a byte-for-byte proof. - 2026-08-28
I-B53
An income statement filed without its cost line does not publish a clean gross margin. The investigation started in the wrong place: mutation testing flagged company-years where the quarters summed to two or three times the full year, and it was recorded as an inflated quarterly series of unknown cause. The opposite was true. The quarterly filings were right and the annual one was malformed, reporting zero cost with gross profit equal to revenue, and the annual revenue figure matched the estimated annual GROSS PROFIT rather than revenue. A hundred and seven company-years publish a gross margin of exactly 100%, which does not exist in an operating company, and the number ships clean, so anyone sorting the market by gross margin gets them at the top. - 2026-08-27
I-B52
The sum of the quarters cannot contradict the financial year. The mutation that prompted this was the least interesting part: while calibrating the threshold, the worst cases came in at 740 times and turned out to be financial years already flagged for a currency-scale problem. The comparison was detecting the same defect through an independent path, and that became a second scale detector which reaches what the first cannot — the last filing of a series has no following comparative, so it was undiagnosable by construction. The flagged population went from 27 to 45. - 2026-08-27
I-B51
Interest on own capital does not vanish from shareholder payouts. A mutation deleted every such entry and no invariant complained, which would have halved the dividend yield of every company that pays it, silently. This is already a declared case in the public challenge, since that instrument is booked as a financial expense rather than a distribution, so anyone summing only the dividend line understates real remuneration. There was a challenge case and no invariant: the challenge proves we can defend that case, the invariant stops it breaking unnoticed. - 2026-08-27
I-B50
The quarterly price-to-earnings ratio equals market cap divided by trailing twelve-month profit. A mutation shifted it by 40% and nothing complained: the annual indicators were recomputed by an earlier invariant, the quarterly series had no equivalent. Same pattern that has now appeared six times in this project, a rule applied to one layer and missing from the neighbouring one. - 2026-08-27
I-B48
A recorded restatement must actually show a divergence. The existing check asked the opposite question, whether every real divergence was recorded; nobody asked whether every record corresponded to a divergence. The asymmetry is easy to miss and the effect is bad in both directions, since the restatement count is a headline number on the home page and an empty row inflates a transparency argument with nothing. - 2026-08-27
I-B47
The version stamped in the database is the version of the code that built it. A mutation replaced the stamp with an invented label and nothing complained. That stamp governs the publication gate and every claim that a user knows which state they audited. A wrong stamp is worse than a missing one, because whoever cites the version ends up citing one that never existed. - 2026-08-27
I-B46
Return on equity and on assets also recompute from the facts, not just the simple ratios. Round one of the mutation audit found that no indicator was recomputed at all and produced the first recomputation invariant; the detection rate then hit 100%, which almost always means the attack set is too easy rather than the system being safe. Round three swapped one company's return on equity for another's: perfect shape, plausible value, normal range, only the owner wrong, and it passed. These two had been left out because they average two financial years, which is more work to reconstruct. More work is not a reason, it is exactly where defects hide, because whoever writes the test also picks the easy path. - 2026-08-27
I-B45
Impossible values are never published: non-positive prices or market caps, or a market cap beyond any plausible order of magnitude. A mutation flipped a price sign and nothing complained. A negative price is not a wrong number but an impossible one, and an impossible value passing means nobody watches that layer. The worst incident on record here was a market capitalisation of 1.67 trillion reais, impossible before it was wrong. - 2026-08-27
I-B44
A flag the rules require must actually be present. Two mutation survivors: erasing a flag from one indicator and a shell-company warning from another went unnoticed. The existing checks verified a flag was CORRECT when present, never that it was PRESENT when due. Flags carry nearly every judgement in this base, so a vanished one returns the number to the world looking clean, which is the worst way to be wrong because nobody reading it suspects. - 2026-08-27
I-B43
Every published indicator recomputes from the facts that produced it. Found by adversarial mutation testing: defects were injected into the data to measure how many the suite caught, and doubling a net margin without touching any fact went completely undetected. The product's central promise, that the published formula is the applied formula, had no invariant at all — it was checked by hand whenever someone thought to look. The drift needs no bad faith: recompute the facts and forget the indicators, which happened in this very project. - 2026-08-26
I-B42
The scale warning travels on the FACT, not only on the indicator. Found by an external audit: the scale-divergence rule suppressed absolute indicators and multiples, but the underlying facts still shipped clean, so a utility published equity of 1.6 million where the following year's comparative says 1.6 billion. Anyone reading the raw facts or the bulk dump got a number that could be a thousandfold wrong with nothing marking it. Marking rather than suppressing is deliberate here: ratios come from these same facts and remain correct, because the scale factor cancels between two accounts of the same vintage, so suppressing would kill good data to remove bad. The general rule it closes: a warning confined to one layer is not a warning, because whoever consumes the layer below never sees it. - 2026-08-26
I-B41
An account the regulator's file reports twice with different values is not published. The archive repeats the same account, same financial year, same column, with contradictory amounts and nothing to tell them apart: same statement group, same dates, same description. One company files profit as both one real and zero on the same line. Until now this resolved by accident on both sides: our ingestion kept the last row read, and the reconstruction script we wrote to AUDIT ourselves kept the first. The two disagreed about the same company, and that is the only reason the conflict surfaced. Arbitrary resolution does not announce itself as arbitrary; it only shows when two of your own parts choose differently, and most systems do not have two parts reading the same source by independent paths. - 2026-08-26
I-B40
Closing equity equals opening equity plus transactions with owners plus comprehensive income plus internal movements. Born from an audit that could not finish: a reviewer built this bridge for one company and stopped halfway, because dividends and other comprehensive income were exposed nowhere in the API, so they could reach a suspicion but not a conclusion. Their finding turned out to be a false alarm, but the hole that prevented them from confirming it was real, and this bridge is precisely the check that catches an equity error. Run across the whole base, 5,347 of 5,361 financial years close. The fourteen that do not are defects in the source, one company filing every closing balance as zero, and those are suppressed rather than published: an internally inconsistent set is worse than an absent one, because whoever checks it concludes the error is ours. - 2026-08-25
TR32
A failing build becomes visible state, not just a log line. Real incident: a change broke one path of the build, the exception was logged, the job retried hourly, the previous database kept serving, and the site looked perfectly normal for sixteen hours. The design protected users, since nothing half-built was ever published, but it protected too well and hid the problem from us as well. The fix is not email: it is publishing the state where someone already looks. The health endpoint is polled by the daily routine and open to anyone, so a silent failure now requires someone to ignore a field that says failed. The test also demands the error itself and a plain-language line explaining what the failure means, because a bare failure count only alarms whoever already knows how to read it. - 2026-08-25
TR29
Wherever an agreement rate is published, the statistical caveat sits in the same block. Raised by the external reviewer, and he was right: this page can produce exactly the impression of certainty it exists to fight. Publishing eighty out of eighty reads as the base has no errors. What it means is only that no divergence was found in those eighty cases, which supports no claim about the population, especially when most of the checks were run by us. The test demands adjacency, not existence: a true caveat alone in a footer is a caveat nobody reads, and serves only to defend us after someone has already misread the number. It is enforced in both languages, since the reader least able to check the rest of the site is the one reading the translation. - 2026-08-25
TR27
The audit sample is a citable artifact, not just a dynamic page. Suggested by the external reviewer and adopted in full: hashing fixed reproducibility, but three years from now the universe will be a different size, and someone who published an audit could not PROVE a given case belonged to that sample. With the sample key and its score the proof is arithmetic and does not depend on us — anyone recomputes the score from the key and the seed. The audit id ties the four parameters that define one run into a short citable label so two different audits cannot be confused. The test also enforces machine-readable quality: prose serves a human, but an API consumer has to decide without interpreting text, so 'this financial year has a scale mismatch, do not use its absolute value' must be a boolean, not a paragraph. - 2026-08-25
TR26
An auditor's sample survives an update to the base. Found against our own interest: running the external reviewer's three seeds on the rebuilt base to send them a case-by-case diff, there was no diff at all — none of the 75 cases recurred. The universe had gone from 63,406 rows to 63,375 and the draw was POSITIONAL, so removing 31 rows shifts every other one. That destroyed the only property this endpoint sells: the auditor publishes a result, the base updates, and nobody can re-check what they claimed. Worse, anyone trying would see different cases and conclude the sample had been hand-picked, which is the accusation the endpoint exists to make impossible. Hashing the key instead of the position gives each case a position of its own, so an update moves only the rows it actually touched. The test simulates the update rather than describing it. - 2026-08-25
I-B39
No fallback flag may cover an entire class of companies. Born from a mistake made the same day the rule was written: revenue growth started using the restated comparative, with a stamped fallback for filings that have none, but the lookup matched the item name used by ordinary companies while financial institutions store it under a different name. Every single bank series fell into the fallback and none used the new rule, while the commit claimed banks were covered. The defect is dangerous because it looks tidy: each number carries a flag, and a flag reads as an explanation. Only the proportion gives it away. An exception that applies to everyone in a class is not an exception, it is the rule failing for that class. - 2026-08-25
I-B37
The revenue-growth denominator is the RESTATED comparative, not the figure as originally filed. Found by reconstructing the published sample against the raw CVM archives, after an external reviewer could audit only a handful of cases because their tooling refuses zip files. The published formula said revenue(year) / revenue(year-1) - 1 without saying WHICH version of the prior year. For one retailer the two readings differ six-fold. The comparative is the right choice for the same reason that governs the rest of the product: whoever opens a year's statement does so the following year, and the comparative shown there is already restated. Mixing the old denominator with the new numerator compares two different accounting vintages. Cases with no comparative fall back to the original figure and are stamped with a flag, so a reader can always tell which base produced the number. - 2026-08-22
I-TR05
Every test in the suite has a description in Portuguese AND in English — the double gate exists because requiring only the translation would let an empty description come back, translated as empty.
Numbers on this page are live. Data version 2026-09-01-supressao-declarada · page generated 2026-09-08 12:05 UTC. If this does not match /saude, you are reading a cached copy.
Home ·
Methodology ·
Transparency ·
Audit it