Dinoer

Running a research campaign

The manifest, the two files a campaign writes, the fields that make the report usable, and how to question a corpus already collected with campagne.py.

campagne.py sits above the browser core: SearXNG discovery, lightweight HTTP collection, escalation to rpa.py/shot.py only for the pages that need a browser, and a report written by a model from the pages collected.

# Git-clone installation
/opt/dinoer/venv/bin/python3 /opt/dinoer/campagne.py --manifeste manifeste.json
# .deb package
dinoer-campaign --manifeste manifeste.json

The manifest

Only id_campagne and cibles are required; every other field has a default. The example documented in campagne.py:

{
  "id_campagne": "concerts-finistere-2026-07-28",
  "cibles": [
    {"type": "query", "valeur": "concerts finistere sud ete"},
    {"type": "url", "valeur": "https://exemple.fr/agenda"},
    {"type": "produit", "valeur": "boulangerie Corentin a Quimper", "max_candidats": 5},
    {"type": "table_reference", "valeur": "mairies du Finistere sud", "cle_thematique": "mairies_finistere_sud"}
  ],
  "max_resultats": 15,
  "max_pages_par_hostname": 3,
  "delai_min_secondes": 4.5,
  "delai_max_secondes": 8.5,
  "revisite_apres_jours": 30
}

Four target types: query (a SearXNG search), url (a fixed page), produit (collect several candidate pages for a named thing and have the model pick the one that is really about it), table_reference (build or reuse a stored list of trusted sites for a subject, then search each of them).

Each target is processed on its own: one that fails does not stop the others, whereas a scenario’s action list stops at the first failure.

Two files, two purposes

operations.jsonl is the operations journal shared by every Dinoer program. A campaign adds one line per attempt, successful or not, tagged intention="dinoer-campagne".

<campaigns directory>/<id_campagne>/collecte.jsonl has one line per page retained, and only those. The report is built from it, and it is the file to open to check that a claim comes from a collected page. The campaigns directory is /var/log/dinoer/campagnes unless DINOER_CAMPAGNES_DIR or the campagnes_dir key of the configuration says otherwise.

Reorder the pages before the report cuts them

By default, the report step puts the pages end to end in the order they were written and cuts at 60,000 characters, without ranking. Two optional fields of the manifest reorder the pages; neither removes any:

{
  "motifs_annee": ["2026"],
  "motifs_mois": ["août", "aout", "/08", "-08-"],
  "sujet_synthese": "concerts and festivals, south Finistère, summer 2026"
}

motifs_annee/motifs_mois is a text filter, without any model: pages that clearly do not mention the requested period go to the end. Put numeric forms such as /08 in the list, because some calendar pages never write the month in letters. sujet_synthese ranks the pages by closeness in meaning to that sentence, with Ollama embeddings. Use both. On the reference campaign, the text filter alone left the official programme of events 27th of the 29 pages it kept, past the cut; the ranking by meaning moved it to 4th.

Ask an open question of a corpus already collected

/opt/dinoer/venv/bin/python3 /opt/dinoer/campagne.py --extraire-cible "<question>" \
  --id-campagne <id> --format-extraction markdown

--extraire-cible accepts any question in natural language, not only a lookup for one fact; --corpus <file> points it at a collecte.jsonl elsewhere. On the territorial campaign of August 2026, 47 calls with an open question returned 17 positive answers. On one page, the narrow question had kept a single date of a five-day festival; the open one returned the whole programme, and four other events from the same text.

To group results that describe the same event on several pages, lib/extraction.py::fusionner_evenements() merges them and keeps every source address. Replayed on the 15 positive extractions of that campaign, it found the same 11 events as the count made by hand. A page that describes several events is cited in each group it belongs to.

The model that writes the report stays off the web

On every call, Dinoer forbids the model its websearch and webfetch tools, whatever the directory the campaign is launched from. Without this, a model once made twelve web searches of its own, invisible in the text it returned. What this covers, and the gap that remains →

Pacing and freshness

delai_min_secondes/delai_max_secondes set a random pause before each SearXNG query and before each page fetched. revisite_apres_jours (30 by default) stops a campaign from fetching again a page it collected recently. The same rules, in full →

The search cache

lib/cache_recherche.py stores in ChromaDB the pages collected for each query. A later query close in meaning is served from it, without a new search or fetch; those pages are marked "source": "cache" in the corpus. --desactiver-cache skips the cache for one run; --purger-cache and --purger-cache-avant-jours N empty it. If ChromaDB or Ollama is unavailable, the cache is simply missed and the campaign goes on.

In short

  • Only id_campagne and cibles are required.
  • Four target types: query, url, produit, table_reference.
  • collecte.jsonl is the file a claim traces back to: open it rather than trust the report alone.
  • Set motifs_annee/motifs_mois and sujet_synthese together so the report does not cut the useful pages.
  • An open --extraire-cible question, page by page, gets more out of a corpus than a narrow one.
  • The model that writes the report is kept off the web, except through bash.