Running a research campaign
The manifest, the two files a campaign writes, the fields that make the report usable, and how to question a corpus already collected with campagne.py.
campagne.py sits above the browser core: SearXNG discovery, lightweight
HTTP collection, escalation to rpa.py/shot.py only for the pages that need
a browser, and a report written by a model from the pages collected.
# Git-clone installation
/opt/dinoer/venv/bin/python3 /opt/dinoer/campagne.py --manifeste manifeste.json
# .deb package
dinoer-campaign --manifeste manifeste.json
The manifest
Only id_campagne and cibles are required; every other field has a
default. The example documented in campagne.py:
{
"id_campagne": "concerts-finistere-2026-07-28",
"cibles": [
{"type": "query", "valeur": "concerts finistere sud ete"},
{"type": "url", "valeur": "https://exemple.fr/agenda"},
{"type": "produit", "valeur": "boulangerie Corentin a Quimper", "max_candidats": 5},
{"type": "table_reference", "valeur": "mairies du Finistere sud", "cle_thematique": "mairies_finistere_sud"}
],
"max_resultats": 15,
"max_pages_par_hostname": 3,
"delai_min_secondes": 4.5,
"delai_max_secondes": 8.5,
"revisite_apres_jours": 30
}
Four target types: query (a SearXNG search), url (a fixed page),
produit (collect several candidate pages for a named thing and have the
model pick the one that is really about it), table_reference (build or
reuse a stored list of trusted sites for a subject, then search each of them).
Each target is processed on its own: one that fails does not stop the others, whereas a scenario’s action list stops at the first failure.
Two files, two purposes
operations.jsonl is the operations journal shared by every Dinoer
program. A campaign adds one line per attempt, successful or not, tagged
intention="dinoer-campagne".
<campaigns directory>/<id_campagne>/collecte.jsonl has one line per page
retained, and only those. The report is built from it, and it is the file to
open to check that a claim comes from a collected page. The campaigns
directory is /var/log/dinoer/campagnes unless DINOER_CAMPAGNES_DIR or the
campagnes_dir key of the configuration says otherwise.
Reorder the pages before the report cuts them
By default, the report step puts the pages end to end in the order they were written and cuts at 60,000 characters, without ranking. Two optional fields of the manifest reorder the pages; neither removes any:
{
"motifs_annee": ["2026"],
"motifs_mois": ["août", "aout", "/08", "-08-"],
"sujet_synthese": "concerts and festivals, south Finistère, summer 2026"
}
motifs_annee/motifs_mois is a text filter, without any model: pages that
clearly do not mention the requested period go to the end. Put numeric forms
such as /08 in the list, because some calendar pages never write the month
in letters. sujet_synthese ranks the pages by closeness in meaning to that
sentence, with Ollama embeddings. Use both. On the reference campaign, the
text filter alone left the official programme of events 27th of the 29 pages
it kept, past the cut; the ranking by meaning moved it to 4th.
Ask an open question of a corpus already collected
/opt/dinoer/venv/bin/python3 /opt/dinoer/campagne.py --extraire-cible "<question>" \
--id-campagne <id> --format-extraction markdown
--extraire-cible accepts any question in natural language, not only a
lookup for one fact; --corpus <file> points it at a collecte.jsonl
elsewhere. On the territorial campaign of August 2026, 47 calls with an open
question returned 17 positive answers. On one page, the narrow question had
kept a single date of a five-day festival; the open one returned the whole
programme, and four other events from the same text.
To group results that describe the same event on several pages,
lib/extraction.py::fusionner_evenements() merges them and keeps every source
address. Replayed on the 15 positive extractions of that campaign, it found
the same 11 events as the count made by hand. A page that describes several
events is cited in each group it belongs to.
The model that writes the report stays off the web
On every call, Dinoer forbids the model its websearch and webfetch tools,
whatever the directory the campaign is launched from. Without this, a model
once made twelve web searches of its own, invisible in the text it returned.
What this covers, and the gap that remains →
Pacing and freshness
delai_min_secondes/delai_max_secondes set a random pause before each
SearXNG query and before each page fetched. revisite_apres_jours (30 by
default) stops a campaign from fetching again a page it collected recently.
The same rules, in full →
The search cache
lib/cache_recherche.py stores in ChromaDB the pages collected for each
query. A later query close in meaning is served from it, without a new search
or fetch; those pages are marked "source": "cache" in the corpus.
--desactiver-cache skips the cache for one run; --purger-cache and
--purger-cache-avant-jours N empty it. If ChromaDB or Ollama is unavailable,
the cache is simply missed and the campaign goes on.
In short
- Only
id_campagneandciblesare required. - Four target types:
query,url,produit,table_reference. collecte.jsonlis the file a claim traces back to: open it rather than trust the report alone.- Set
motifs_annee/motifs_moisandsujet_synthesetogether so the report does not cut the useful pages. - An open
--extraire-ciblequestion, page by page, gets more out of a corpus than a narrow one. - The model that writes the report is kept off the web, except through
bash.