Atlas · MACRO NOTE
Published 2026-08-10

Database Reconciliation and Search Engine Optimization Audits in the Atlas Research Notebook

Reconciliation of the Atlas Search Console metadata and database registries has established a verified indexing footprint of 45,017 indexable stocks and stub routes while certifying 56,892 active PDF documents on the public research surface.

Table of Contents

Google Search Console Indexation Audit

A comprehensive reconciliation of the Google Search Console export dated 2026-08-09, reflecting data as of 2026-08-07, reveals exactly 69,238 known URLs associated with the research domain. Within this catalog, 34,170 URLs are indexed and 35,068 URLs are not indexed. An inventory of the /var/www/atlas/stocks directory identifies 62,332 total folders, comprising 62,317 symbol and stub pages. This collection divides into 45,017 indexable pages (27,092 clean tickers and 17,925 indexable venue pages) and 17,300 noindexed pages. The noindexed pages represent 27.76% of the stock directory fleet. Within this noindexed subset, 98.27% are venue-prefixed stubs carrying the <meta name="robots" content="noindex,follow"> tag and canonical links pointing back to clean ticker parents.

The sitemap configuration index at sitemap-index.xml submits exactly 45,053 URLs across seven tiered sitemap files: hot1000-sitemap.xml, stocks-t1-indexed.xml, stocks-t2-populated.xml, stocks-t2-populated-2.xml, stocks-t2-populated-3.xml, stocks-t3-identity.xml, and stocks-t3-identity-2.xml. Audit logs for atlas_symbol_browse_pages.py verify that browse links only expose sitemap-listed indexable pages, generating 27,187 unique stock links while explicitly excluding venue stubs, with only 28 stubs linked out of 17,001. The approximately 24,700 unsubmitted but indexed URLs in the Search Console originate from a legacy stocks/discovery-sitemap.xml dump of 35,200 URLs and historical deep link crawling.

Google Search Console records daily impression recovery starting in late July and early August, peaking at approximately 9,500 impressions on August 2, 2026, as search crawlers re-evaluate the 45,025 submitted sitemap URLs. Over the three-month period from May 11 to August 6, 2026, totaling 88 days, the research domain recorded 65,449 impressions, 51 clicks, an average click-through rate of 0.0779%, and an average position of 40.83. During the latest seven-day period, traffic expanded to 44,091 impressions and 35 clicks, representing a click-through rate of 0.0794% and an average position of 47.09, compared to 4,955 impressions and 3 clicks in the preceding week. Mobile search achieved a 0.45% click-through rate (26 clicks from 5,821 impressions) compared to a desktop click-through rate of 0.04% (25 clicks from 59,355 impressions). Formatting title metadata to avoid synthetic query terms remains a primary target to improve conversion rates for organic search traffic targeting the Atlas Pro repository.

Database Reconciliation and Document Validation

The active document library contains 56,892 verified files, comprising 28,370 exact dark and light theme listing pairs alongside 76 legacy ticker pairs. A current published writer roster of 35,666 entries intersects with the validated files at 28,365 pairs, leaving a backlog of 7,301 items and 76 verified symbols outside the active roster. Within this backlog, 6,268 missing pairs lack a public fields schema key and 748 pairs fail canonical naming validations. Earlier diagnoses of a broader gap were corrected after verifying that compiler file counts include ticker, slug, and case aliases rather than representing unique markets. Analytical discrepancies occurred because a compiler script marked contexts as validated on both success and failure branches, storing synthetic dossiers that lacked direct database grounding. A separate rendering script fell back to bulk-rewritten timestamps, causing the 45-minute selector timer to repeatedly process the same ten unchanged symbols: 8306, DIAL.N0000, ICT, IHC, KFH, MAYBANK, MTNN, RHO, SBER, and SCOM.

Audit metrics show 147,413 files in the broader environment, comprising 89,989 backups, 56,892 active library PDFs, and 532 historical files, with 94,030 unique content hashes and 53,375 duplicate files. To enforce data integrity, four source-backed pairs were validated and released: NASDAQ:COST containing 10 pages, 3,466 source words, and 16 sources; NYSE:GE containing 11 pages, 3,032 source words, and 14 sources; NYSE:MCD containing 13 pages, 3,647 source words, and 19 sources; and NYSE:XOM containing 10 pages, 2,863 source words, and 15 sources. These filings extract cleanly across text layers, and their SHA-256 values match the central registry. Conversely, five candidate listings were held to prevent ticker collisions: LSE:BA. (where BAE Systems and Boeing collided), SGX:F&N (Fraser & Neave and Fabrinet), LSE:SN. (Smith & Nephew and SharkNinja), and SAP and SILVER due to low content volume. This rigorous filtering keeps Atlas symbol coverage anchored in clean disclosures.

Streaming Protocol Performance and Knowledge Retrieval

The interface located at the public ask route has transitioned to a single terminal script, consolidating legacy scripts and styles. Real-time streaming is enabled through a streaming query endpoint that disables proxy buffering and transmits NDJSON event sequences. Initial benchmarks from a route audit recorded the first byte at 1.332 seconds, the first visible token at 1.826 seconds, 405 individual token events, and complete query resolution at 15.314 seconds. Search queries are verified against a public receipt of truth at https://atlas.freedomcore.io/data/atlas-truth.json, which certifies 43,303 written research briefs, 13,811 validated briefs, and 28,889 briefs with verified intelligence. The eight-day public freshness subset stands at 13,810 written briefs. This validation architecture allows users to search the Atlas notes with precise corpus references. OpenBB documentation from the OpenBB docs and DeFi schemas from the DefiLlama API docs support these data-source methodologies, structuring the OpenBB DefiLlama RSS market spine for accurate retrieval.

Economic Calendar and Market Integration

Macroeconomic indicators and public economic calendars are updated via the Atlas Macro Radar, integrating economic calendar research with data from the BLS, BEA, ONS, ECB, and Federal Reserve. To protect query performance, non-macro questions no longer query FRED series. This change removes the 5 to 8 seconds of latency previously caused by fetching 14 FRED series during standard queries. Static page building resolves exchange venue aliases dynamically to support cross asset market intelligence.

Analytical Caveats and Development Boundaries

A gap of 16,546 listing pairs remains missing from the published catalog. The published catalog contains 44,916 exact listing identities, while the active library contains 56,892 READY files (28,370 pairs and 76 legacy pairs). The missing backlog contains 7,296 pairs addressable by ticker mapping and 9,250 pairs across 6,278 collision groups held for deduplication. A listing-key-native writer is under development to resolve these holds. Incremental compilation is currently stopped (paused on August 7, 2026, at 397 of 16,764 items). The translation engine is not yet live; all 56,892 active PDFs are currently English-only. Crawlers are blocked on 1,333 URLs due to robots.txt rules. Disk capacity is 91% used (42G free), which restricts bulk-rendering operations pending a storage preflight guard. All database modifications and schema releases remain subject to operator review.

Browse the Atlas research notebook

FreedomCore Atlas Research →