Two official languages mean two addresses for every idea, and a search engine applies one crawl allowance to both. This is about the step nobody reports on: how an address becomes known, what drains that allowance on a document-heavy organization, and how much of it the Semalt panel's Indexing Hub lets you observe.

Ask a communications lead in the National Capital Region for a page count and the CMS supplies one to the digit. Ask what a search engine holds and there is no answer. On most sites here the second figure trails the first by a wide margin, and nothing raised a warning.

The cause is structural. Whatever an association, a federal supplier, a university faculty or a consulting firm publishes appears twice, English and French, before anyone commissions new material. Add a PDF publications library, a member directory, tender and program pages tied to the April-to-March cycle, and an archive nobody has pruned since 2011, and the address count outgrows the editorial plan several times over.

What this covers. Ranking factors and writing quality belong elsewhere. The subject here is earlier: whether a bot ever hears about an address, whether it spends a request on it, and how you watch both.
Starting point · Three steps, one word

Published is not the same condition as found

Between publication and a search result sit three steps, each capable of failing on its own terms. An engine has to hear that the address exists. A bot has to request it and receive something usable. Last, the engine weighs whether to store what came back. Compressing all three into the word "indexing" costs teams quarters, because the remedies do not carry across: editing a page achieves nothing while no crawler has asked for it, and announcing it again achieves nothing once it has been fetched and rejected on merit.

  • Discovery is a matter of notification. If the French version of a policy brief hangs off a toggle no crawler will operate, it was never announced.
  • Crawling is a matter of capacity. A bot allots limited effort to your domain each day, and directory filters absorb every bit of it.
  • Retention is a matter of judgement. A two-line program stub, or an HTML shell holding nothing but a download link, loses that judgement.
  • Each step leaves its own evidence. An empty visit log points to notification; a logged visit with a clean response and no presence in results points elsewhere.

An example fixes the distinction. An association moves its bilingual publications library, landing 1,600 documents on new paths. Ten weeks on, about 500 can be found. If the log shows no bot calling on the other 1,100, the failure is notification and a push resolves it. If all were fetched normally, pushing changes nothing.

Locate the step first. The opening question is not how to get pages indexed but which step they stopped at — a matter of record, not of opinion.
Mechanics · Where requests land

Crawl allowance, and what drains it

A crawl allowance is not issued, requested or appealed. It falls out of two continuous measurements: how quickly and dependably the server answers, and how much apparent demand the domain has earned. A repository that stalls under load in late March is then treated cautiously for weeks, because crawlers throttle themselves rather than topple a struggling host.

Most requests land on addresses no editor asked for. Software mints them — a directory, a document viewer, a language switcher, an events calendar — and here every generator emits its output twice.

Source of the addresses How it shows up locally Order of magnitude Who owns the fix
The second language tree Every notice, program and biography duplicated under /fr/ Doubles the count on day one Nobody — planned around, never removed
Document viewers One PDF as a viewer page, a file and a fragment Three to five per document Development: one canonical route
Member directories Filters for province, sector, class and language, all answering Tens of thousands of states Development: robots rules, canonical tags
Dated pages and unpruned archives Closed opportunities, last year's conference, a decade of pagination A fresh cohort each April Communications: retire and consolidate

Attach figures and the shape is plain. Nine thousand genuine pages become eighteen thousand addresses the moment French exists, four thousand documents behind a viewer contribute a multiple of their count, and Tuesday's consultation page queues behind all of it.

9,000
genuine content pages
18,000
once French exists
60,000+
reachable directory states
400
pages with proven demand
Removal beats submission. Taking fifty thousand pointless addresses out of a crawler's way helps more than any volume of announcements, and costs configuration time rather than a line in next year's estimates.
Patterns · Four shapes in this market

Large by mandate rather than by output

The bodies most affected are seldom those publishing most often. Their page count accumulated as a by-product of obligation.

Associations

Members, positions, proceedings

A directory of several thousand members, a policy library and a decade of proceedings, in both languages.

  • Filter states outnumber members
  • Old event microsites never taken down
Suppliers

Capability pages on a tender cycle

Firms bidding on federal work keep capability and past-performance pages that need a credible French version to count.

  • Content expires with the fiscal year
  • Closed opportunities stay published
Universities

Faculties and repositories

Program descriptions, staff profiles, research outputs and a repository, each faculty on its own rhythm, none deleting.

  • Course archives kept for accreditation
  • Items repeated across faculties
Professional services

Practice pages and commentary

Bilingual practice descriptions, partner biographies and years of commentary aimed at procurement reviewers.

  • Every biography exists twice
  • Commentary ages but stays up

Each shape fails in the same direction. Whatever matters this quarter — a live opportunity, a new program, a submission on a bill before committee — is the newest and least linked thing on the site, while material that stopped mattering in 2014 sits in every archive listing.

Timing complicates it. Because demand follows the fiscal year, comparing one month against the last tells you little. The defensible baseline is the same point in the previous cycle: how much of last April's material was picked up by the end of May.

Notification · The sitemap as an instrument

A sitemap is a declaration, not paperwork

On most sites the sitemap is a file some plugin writes and nobody opens. On a bilingual site with a document library it is the one dependable way to state what exists and what moved, and its construction decides whether the reply is readable.

The Hub accepts one as an upload or as an address, then walks it recursively through three levels: an index pointing at further indexes, which point at the files holding addresses. A single job takes in as many as 1,000 sitemaps, so splitting the tree is realistic at this size.

Structure · Three levels

A tree that survives a full fiscal cycle

Divided by language first and by rate of change second, so a stalled branch identifies itself.

no extra charge
  • A single file at the top. One index that the engine and the Hub are both pointed at, naming the section indexes below.
  • Language decides the second level. English and French each get an index, so "is the site indexed" becomes two answers, not an average.
  • Rate of change decides the third. Stable pages, the publications library, dated material and the archive, each kept apart.
  • Dates have to be honest. A lastmod that shifts on every nightly rebuild teaches a crawler to disregard your dates — a hazard where a committee edit and a deployment look alike to the CMS.
  • Only live canonical addresses are listed. Nothing redirected, nothing missing, nothing carrying a noindex, nothing pointing its canonical elsewhere.

Assembled that way, the file becomes a gauge. If the French publications index reports eighty announced and twenty reached while its English twin reports eighty and seventy-four, the fault is narrowed to one branch.

Stable

Rarely edited

Mandate, governance, contact details, program overviews and standing capability pages.

  • Dates that hold still
  • Announced again only after a real change
Dated

Written to expire

Open opportunities, current-cycle programs, consultations and conferences.

  • Regenerated from the live record
  • Dropped on the closing date
Two moving, twenty waiting. The module runs two sitemap jobs at once and holds twenty more in line. Nothing is discarded, but a migration fired off as forty submissions becomes impossible to trace.
Instrument · Indexing Hub

Four numbers that bound the module

Within the panel, this module sits beside the campaign, analytics and rank-tracking sections and handles everything ahead of a ranking; the tour of the rebuilt Semalt workspace places it there deliberately. Four figures define what it accepts, and none move on request.

Capacity · Build the schedule on these

Daily rate, batch size, depth, concurrency

Any submission plan for a large bilingual site has to fit inside these four constraints.

part of the panel
  • One thousand URLs each day, per account. The tracker's working rate, and unused capacity does not accumulate.
  • Ten thousand URLs in a single batch. This governs how much you hand over at once, not how fast it clears: a full batch is ten days of work.
  • Sitemaps parsed three levels down. Supply a file or an address; nested indexes are followed automatically, to 1,000 sitemaps per job.
  • Two jobs active, twenty in line. Concurrency does not vary, so a bilingual migration runs in sequence, in the order you set.
1,000
URLs each day
10,000
URLs in one batch
3
levels parsed
1,000
sitemaps inside a job

Announcements travel over IndexNow, which notifies participating crawlers — GoogleBot and BingBot among them — that something at a location has appeared or changed. Instead of hoping a bot revisits a corner of the repository it last saw in February, you raise your hand the day the change lands — which matters most for short-lived material, such as a consultation open six weeks.

Say this before anyone plans around it. Submitting a URL is not the same as getting it indexed. An announcement conveys one fact: something at this location changed. Whether a fetch follows, and whether anything is stored, rests with the engine and turns on what the page holds. No vendor here, Semalt included, can sell indexation; one that offers it is promising an outcome it does not govern.

The middle step, though, becomes provable. Every address keeps an entry: which bot appeared, when, what status it got back, and the detail behind any failure. Three counters sit above that log — announced, reached, failed.

2
sitemap jobs at once
20
further jobs in line
3
counters updating live
Observation · Reading the result

What a finished batch is telling you

A completed batch leaves three totals and a log entry per address. The totals mean something only against each other; the log settles whose problem this is — the developers', the server's or the editors'.

Pattern in the totals Most likely explanation Where to look
Many announced, few reached, few failures The notice registered; no bot has come by Internal links; orphans wait longest
Reached nearly matches announced, results flat Fetching is healthy; pages were examined and passed over Thin stubs and PDF landing pages
Failures bunched in the French branch A routing or template defect on one side The error text, then that branch's file
Failures spread evenly and thinly The server ran short mid-crawl Response times at those moments
Clean responses on documents that never surface The engine looked and declined them Nothing here; this becomes a content question

That last line is why per-address logging earns its keep. A recorded normal response proves the bot arrived and the server behaved; about storage it says nothing. Holding the two apart turns "traffic dropped" into a task with an owner.

Take a reading before you start. Record how many addresses in each language branch already show a logged visit. Every later claim is measured from there; without it, two months of work cannot be defended.

Check it weekly rather than daily, and read the text behind failures instead of the count: sixty errors sharing one status code under one path prefix is a defect report.

Complication · Two trees, one calendar

The bilingual duplicate that cannot be consolidated

French is not a discretionary second market here. For an association with members across the country, for anyone bidding on federal work, for any organization serving Gatineau, it is a condition of operating. The usual advice about duplicate content is therefore useless — an obligation cannot be merged away — and what remains is making the two trees legible as two.

Arrangement How it reads to an engine The price Practical remedy
French produced by machine translation A weaker duplicate of everything Half the daily rate on pages nothing retains Rewrite what carries demand; delete the rest
Language chosen by parameter or cookie One location answering differently each visit The French version is invisible Separate paths per language, cross-declared
A PDF competing with its landing page Two locations answering one question Requests divided, neither consolidated Choose one; the HTML page usually deserves it
Dated pages left up past closing A large body of stale but valid pages Allowance spent on the previous fiscal year Retire automatically, keyed to the closing date

Choosing what goes first is a question of evidence. The Search Console views in the same panel report which pages draw clicks and impressions, with a keyword dynamics view following what moves into and out of the top ten. Because both languages carry equal weight here, it is the language split of that data that does the analytical work, never the country filter. Exports run to 10,000 rows in CSV or JSON, and Stream, the assistant attached to My SEO, accepts a URL list in bulk.

Then there is the constraint no software removes. Publication passes through committee, official-languages review and, in government-adjacent work, legal sign-off. Advice assuming a page can be corrected by the end of the day is wrong here, which argues for queuing the structural work early.

My SEO · Around the approval queue

What proceeds while sign-off is pending

Parameter handling and retirement rules need developer time. Keyword and link work does not.

149 / 500 USD monthly · per domain
  • Campaign work proceeds in parallel. AutoSEO, at 149 USD monthly for one domain, covers keyword discovery, prioritization and link building without entering the release process.
  • Manual control where review culture expects it. FullSEO, at 500 USD, adds hand-picked keywords with automatic fallback, placement targeted by domain rating, and human review before on-site changes ship.
  • Four to eight weeks before anything shows. The usual interval to first measurable movement, and a reason to start a cycle early.

Questions that come up

Does pushing a URL through the module get it into the index?

It does not, and a vendor claiming otherwise is overstating. The push informs participating crawlers that something at a location is new or altered. Retrieval, and the choice to keep what was retrieved, stay with the engine.

Should French be handled as a separate job?

Yes, for diagnostic reasons more than technical ones. With an index per language the counters report on each branch separately, so a stall shows up in the totals rather than half a year later, through a complaint from a member in Gatineau.

Do PDFs belong in a sitemap?

Documents that answer a question by themselves — a standard, a guideline, an annual report looked up by name — can be listed. Those sitting behind a proper landing page should not be; list the page, or a crawler sees two locations for one piece of content.

Why does only a thousand a day move if ten thousand can be handed over?

They describe different things: one is the size of a hand-off, the other the rate of processing. A full batch schedules ten days of work, so build the calendar on the daily figure.

Our opportunities close in six weeks. Worth announcing them?

Announce them the day they publish, since six weeks is short next to an unassisted crawl interval. The lasting value sits above them, in capability and sector pages that persist across cycles and stay current because of the opportunities underneath.

Arithmetic · Forty thousand addresses

The closing calculation, and a first week

Take a plausible case. An association replatforms: 9,000 content pages per language, 4,000 documents, a member directory and fifteen years of proceedings. Everything moves to new paths, leaving roughly 40,000 addresses to be rediscovered.

At a thousand a day that is forty days of processing — arithmetic rather than a defect in the tool, worth stating before somebody promises the board a recovery inside a month. Four batches hold the volume; the daily rate governs the calendar.

40,000
addresses after the move
40
days at the daily rate
4
batches of ten thousand
10
days for everything earning

Forty days only hurts if the material that matters lands on day thirty-eight. Sequenced by measured value across both languages, everything earning is through within two weeks.

Days Contents Volume Why in this position
1–4 Mandate, membership, capability and contact, both languages 3,600 Looked up by name; the doorway to the rest
5–10 Current-cycle programs, consultations, live opportunities 5,400 Dated; a week lost cannot be recovered
11–22 Publications with recorded impressions, and their landing pages 11,000 Proven demand ahead of anything speculative
23–40 Directory states, tag listings, pre-2015 proceedings 20,000 Almost no measured value; most wants consolidating

That bottom row deserves scrutiny before it swallows three weeks of capacity. Where twenty thousand addresses drew no impressions in either language all year, the answer is consolidation or deletion, not announcement.

Begin narrowly, and begin this week. Cut the sitemap back to live canonical addresses, divide it by language and then by rate of change, export the pages that already draw impressions, and work that list at the rate the module permits. To keep indexing, analytics and campaign work in one place, open the Semalt dashboard and add your property. What an outside team takes over is set out on our services page, and further reading is on the blog.

The caution once more. None of this obliges an engine to retain a single page. Sitemaps, batches and announcements remove every excuse for a location going unnoticed, in either language. Retention is decided by the page, and that decision is neither yours nor ours.