← Back to home PT
Metodologia aberta

Sector classification — 10 macro-sectors

Sector classification — 10 macro-sectors

B3 has hundreds of companies, and the CVM registry classifies them into 70 activity sectors (SETOR_ATIV, the registry's activity-sector field) — too granular and too noisy to compare sectors, and full of labels like "Emp. Adm. Part. - Energia Elétrica" (holdings) and "Securitização de Recebíveis" (receivables securitization vehicles). For the rankings and the sector view, we fold those 70 into 10 standard macro-sectors (GICS/B3 line):

Financials · Oil & Gas · Basic Materials · Industrials · Consumer Cyclical · Consumer Defensive · Utilities · Healthcare · Tech & Communication · Holdings/Diversified.

How (auditable): the mapping is by keyword over the registry's SETOR_ATIV (robust to the abbreviations CVM uses), with an explicit rule and a defined order — the first that matches wins. It is open in dados_b3/setores.py. Every company carries its source sector, so you can check the classification one by one.

Declared curation rules: - Holdings of an end-sector ("Emp. Adm. Part. - Energia Elétrica") fall into the end-sector (Energy → Utilities). Only the "Sem Setor Principal" (no principal sector) ones become Holdings/Diversified. - Civil construction goes into Consumer Cyclical (following B3: homebuilders are consumer cyclical), not into a separate real-estate sector. - Petrochemicals go into Basic Materials (chemicals), separate from Oil & Gas (exploration and production).

A sector that shows up in the registry and matches no rule is left unclassified on purpose — the absence fails the coverage test, instead of silently falling into the wrong bucket. It is the same discipline as the rest of the database: a number (or a label) that did not pass is left out, not guessed.

Sector total-return index (the "sector race")

This index was pulled offline and rebuilt on 2026-08-15. What was published before that date was wrong. The section "The 2026-08-15 error" at the end of this page tells what happened, why the tests missed it, and what now exists so it cannot happen again. It is here on purpose: a product that promises auditability has to leave its own mistakes auditable.

For each macro-sector, a weekly index rebased to 100, market-cap weighted and with dividends reinvested (total return — not price alone, which understates a high-dividend sector).

  • Week = last trading session of each ISO week. The series starts at the first week with real coverage and runs to today; the window (1 / 3 / 5 years / max) is chosen on the page.

"Real coverage" has two floors, and both were learned the hard way. The first is the market's: at least 30 companies with a price and a weight in the week. The second is each sector's: at least 3 companies — and it holds for the whole series, not just the start, because an average of two companies is not a sector.

Why the series does not start in 2010, even though prices go back to 2010: the weight is market cap, which needs a published filing + a price + a share count. Before ~2021 that full set does not exist for enough companies. An earlier version of this page extended the series to 2010 anyway, and the result was 364 of 865 weeks frozen at exactly 100.0: a week with no members returned a factor of 1.0, absence of data became "the market did not move", and the index then caught up in jumps. The cumulative figure that came out of it was fiction. We prefer a shorter, true series to a long, crooked one — and an invariant now blocks publication if any week comes out frozen. - Sector weekly factor = the average of the total returns of the sector's stocks, weighted by market cap. Only members with a price in the week and in the previous one are included (a stock that IPO'd later joins when it listed — it does not shorten the sector). - Stock total return in the week = (end_price × corporate-action multiplier + dividends with ex-date in the week, for the main ticker's class, without a flag) ÷ start_price.

That multiplier is the fix for the 2026-08-15 error. COTAHIST publishes the price as traded, with no retroactive adjustment: in a 1:4 split the price drops 75% in one session and the holder has lost nothing. We now ingest B3's series of splits, reverse splits and bonus issues (GetListedSupplementCompany) and verify every event against the price drop observed in COTAHIST itself. Three outcomes, all of them recorded:

outcome what we do
the declared multiplier explains the observed drop apply it
the event is too small for price to tell apart from an ordinary session (a ~10% bonus issue) apply it, marked declared by the source, not confirmed by us
the source declares it and the price contradicts it neither apply nor ignore: the stock leaves that week's calculation

Same-day events are composed before the verdict (B3 records "split 100x and reverse-split 0.01x" as two rows — together they make 1.0). And when B3 omits an event, a second independent source finds it: the share count the company declares in its filings. On its own that source is useless (it cannot tell a reverse split from a share issuance), but requiring both to agree — the share count moved by a factor and the price jumped by the inverse in a single session — surfaced 13 events absent from B3's own registry, among them the reverse splits of OIBR3 and SEQL3, which alone were worth +669% and +1,779% jumps in a week. - Weight = annual market_cap (the same one already used in the multiples, with the scale corrected), from the year ≤ the week's year. - "Market (B3)" = the same calculation over ALL covered companies. It is NOT the official Ibovespa — we do not use the index's series; we build our own, with the same yardstick, to compare apples to apples (the official index mixes in another methodology). It is the honest benchmark of our own universe.

Source 100% ours (COTAHIST + B3 dividends + CVM); no external index series. It is a descriptive map — "this is what happened" — and not a decision rule: a sector that went up is not a buy signal. See multiplos.md and dividendos.md.

Declared limitations (read before drawing a conclusion)

1. Survivorship bias. This is the most serious limitation of this index, and it is important that it be written here and not hidden. The index universe is the companies active today in the CVM registry. A company that went bankrupt, was delisted or went private over the period is not in the series — not even for the stretch in which it was still trading. Since what disappears from the exchange tends to disappear after falling, the index of a sector that lost companies is optimistic: it shows the return of those who survived, not the return of someone who invested in the whole sector at the start. The effect is larger in sectors that went through bankruptcy reorganization and consolidation. Fixing this would require rebuilding the historical point-in-time universe (who was listed in each week), with the prices of the delisted ones — we do not have that series, and that is why the limitation is declared rather than estimated.

2. The aggregate hides dispersion. "Oil & Gas" is few companies, and market-cap weighting makes the largest dominate the sector: the number is, in practice, close to that company's return. A large, heterogeneous sector (Consumer Cyclical), on the other hand, has retail falling and education rising inside the same line. That is why the page shows the number of companies in each sector next to its name — an average of two is one thing, an average of seventy is another.

3. Scale. In the chart, the linear scale compares levels; the logarithmic one compares rates — on it, the same slope means the same percentage return, and a sector that multiplied by five stops flattening all the others. Use log to compare trajectories; linear to see the size of the final result.

4. Events we cannot quantify. Five groups of corporate actions in the published window are declared by B3 but contradicted by the price, and a handful of spin-offs (where the holder receives shares of another company) have no possible adjustment without pricing the asset received. In those cases the stock leaves that week. This introduces a small, known bias: the week loses a member. It beats both alternatives — publishing a −59% that never happened, or applying a factor the price itself contradicts.

The 2026-08-15 error

This section exists because the product sells auditability. Hiding our own mistake would be selling the opposite of what we deliver.

What was published, and was wrong. Until 2026-08-15 this page showed a sector index "since 2010" with the market at +518%. None of those numbers were real. Three defects were stacked, and the third only surfaced because we chased the first:

  1. Phantom coverage. 364 of 865 weeks sat at exactly 100.0, because before ~2021 almost no company had a market cap and a week with no members returned a factor of 1.0. The index went flat and then caught up in jumps.
  2. Untreated corporate actions. COTAHIST is not split-adjusted. Oil & Gas showed −89.6% in a single week; "Market" itself showed −42.4%. For calibration: in the week of the COVID crash, the worst-hit sector did −17.5% — and that is a real move.
  3. An impossible market cap. Investigating the weight behind that −42.4%, we found AZUL at R$ 1.67 quadrillion: the 54.7 trillion shares from the 2026-03-31 balance sheet multiplied by a price from 04-20 — already after the 1-for-150,000 reverse split of the 17th. A company worth more than a hundred times the country's GDP went through the entire database without tripping anything.

The patch that did not work, and why that matters. The first attempt was an amplitude filter: drop any stock whose weekly return fell outside −50%/+100%. It failed, and the failure is instructive — amplitude cannot tell a split from a real crash. Six series still carried artifacts and Oil & Gas went from wrong-too-low to wrong-too-high (+1,429%). A statistical filter is no substitute for the data that is missing. That is when the index went offline, and it stayed off until the data existed.

Why the tests missed it. The invariants checked identities (coverage, recomputation, absence of nulls) but none asked "is this number possible?". A large cumulative return trips no alarm — large is what you expect from a long-horizon index. And the cross-check that did exist compared the "since 2010" series with the "since 2021" one and matched; it matched because it only looked at the last five years, i.e. it checked the part that did not have the defect.

What now exists. Four new invariants, every one of them written from a defect that actually happened:

  • no frozen week to the cent, in any sector (previously: up to 2%);
  • no sector week beyond ±35% — the yardstick comes from the data itself (−17.5% for the worst sector in the COVID week);
  • no market cap larger than Brazil's GDP — crude on purpose; it would have caught AZUL on the first build;
  • the multiple's price inside the window after the filing was published.

That last one came as a by-product and is the most serious finding of all: while investigating AZUL's market cap, we discovered the point-in-time price lookup was data >= dt_receb with no upper bound. When the current ticker had no price back then (a company that changed its code, migrated share class, or only started trading later), the query fell forward until it found a price — and EMBRAER's 2010 P/E came out computed with a 2025 price. Fifteen years of look-ahead, in 22% of the multiples rows, silently. The session now has to fall within 45 days of publication; if it does not, the multiple is not computed. The root fix — ingesting prices by ISIN instead of by ticker, as we already do for REITs — is queued.

What it cost in coverage. The published series shrank from "since 2010" (fiction) to since 2021 (real), and part of the historical multiples ceased to exist. That is the price of trading a wrong number for a declared absence, and in this house that price gets paid.