Checking, converting and comparing

PLATO tools works on data in PLATO’s shape. It checks a file for mistakes, converts it to another format, and shows whether a new version of a published dataset has kept what the earlier one said, and it prepares a dataset for publishing. It runs in your browser: your files are never uploaded, and it works for datasets of any size your disk can hold. It also runs from the command line, for many files at a time, with the same checks and the same reports.

To use it, open the page and drop a file on it (or choose one), then press Check, choose a format and press Convert, press Compare with the earlier version, or choose what to prepare for publishing and press Prepare.

What it reads and writes

Format

Read

Write

The spreadsheet tables: the ten CSV files, a zip of them, or the temPlato workbook

yes

yes, as a zip of the ten CSV files

PLATO JSON, place-centric or attestation-centric, and JSON Lines

yes

yes (place-centric)

Linked data: N-Triples, N-Quads or Turtle

yes

N-Triples

Linked Places Format, version 1

yes

yes

Annotations from Recogito

yes

no

A gzipped file (ending .gz) is read as it is.

Checking

Check holds a file to PLATO’s own definitions. A JSON document is checked against PLATO’s JSON Schemas; spreadsheet tables against their table definitions, and against the rules those definitions cannot state, such as that a relation names a place or something else, one or the other; linked data against the terms the ontology defines. It also finds two mistakes no schema can: a route or network that, through its members, is a member of itself; and an identity match written under one place that names another place as its subject.

The report comes in three parts:

  • Problems must be put right before the data is valid PLATO. Each is named in plain words, with where it is: the sheet, row and column, or the place and the key.

  • Warnings are worth a look, but the data can still be used.

  • Would not be carried over lists what PLATO JSON has no place for, should you convert the file.

A file that stops part-way, or is damaged, is reported as a problem in the file, with what could be read before it.

Converting

PLATO JSON and linked data hold everything PLATO can say, and converting between them loses nothing. The spreadsheet tables and Linked Places Format hold less, by design: see what the spreadsheets cannot say. Two rules govern what PLATO tools does about that.

Nothing is left out without saying so. The report names everything the target format has no place for, in words (“The pronunciation of a name…: Linked Places Format has no place for this, so it is left out”).

Nothing is written that would mislead. Where leaving a detail out would change what a statement says, the whole statement is left out, and the report says so:

PLATO JSON, linked data

Spreadsheet tables

Linked Places Format

A denial: the source says something was not so

kept

kept, one thing denied per row

left out

A statement that a later one withdraws or replaces

kept, with what withdraws it

left out

left out

A computed value, worked out by software

kept, marked as computed

left out

left out

A figure from a statistical table

kept

left out

left out

Identity matches made together, in one statement

kept

left out

left out

How firmly the source says it (stance)

kept

kept

the statement is kept; its stance is reported as left out

So a file in the tables or in Linked Places Format shows the dataset as it stands now. It never shows a withdrawn statement as current, a denied market as a market, or one figure from a table as a fact about the whole place.

Web addresses for your identifiers. PLATO fixes how the spreadsheet tables’ identifiers become web addresses, so that every tool gives the same ones. The base is the one given for the conversion, if any (on the page, “Web address for your identifiers”), and otherwise the base_uri of the about sheet, which is where it belongs; a / is added to it unless it already ends in / or #. A place’s address is the base, then place/, then its place_id (https://w3id.org/my-project/place/bristol), and a source’s is the base, then source/, then its source_id. In both, every character other than a letter, a digit or one of - . _ ~ is percent-encoded (written as a code such as %20), so keep identifiers to those characters. The dataset’s own address is the about sheet’s dataset_uri, or the base (with its /) if that is empty. Without any base, PLATO tools uses a stand-in, https://example.org/my-dataset/, which is not a permanent address. Converting to the tables, an address is kept only if reading the tables back would give the same one; otherwise the report says it is lost.

Comparing two versions

Once a dataset says it is published, its attestations are only ever added to: none is deleted or changed, and a correction is a new attestation that withdraws or replaces the old one, which stays. Why, and what that makes possible, is explained under citing a place in a gazetteer that changes.

Compare with the earlier version shows that a new version has kept to this. Choose the new version first, then press the button and choose the earlier one. The two may be in different formats: what is compared is what each attestation says, not how the file writes it.

Problems are what breaks the rule:

  • an attestation of the earlier version that is missing from the later one;

  • an attestation that says something different in the later one;

  • an attestation that says the same but has lost its web address, or has been given another: later statements point to it by that address;

  • a name, location, date, type, property or citation, with a web address of its own, that attestations point to, and that the later version describes differently, or no longer describes at all though attestations still point to it. It is part of what those attestations say, so changing it changes them.

For the first few attestations, names and other facets that changed, the report shows what changed: what one version says and the other does not. Each problem says how to put it right.

Warnings do not break the rule:

  • a place or a source described differently, since those may be corrected;

  • an identity match removed or changed;

  • an attestation added without the date it was made (created), without which the dataset’s state at an earlier moment cannot be worked out;

  • a later version that is not marked published, gives the same version as the earlier one, or names another as the version before it.

Before the earlier version is published the rule does not yet apply, so everything is listed as a warning, and the report says why.

Give your attestations web addresses. An attestation without an address of its own (in JSON, an @id) can only be found by what it says. If one goes missing, the check cannot tell whether it was deleted or changed, and no later attestation can withdraw or replace it, since there is nothing to point to. The report counts such attestations. PLATO recommends that every attestation in a published dataset has a permanent address. The spreadsheet tables have no column for one, so this applies to data published as JSON or linked data.

A comparison that could not read the whole of either version does not pass, and nor does one whose earlier version holds no attestations: in both cases nothing, or not everything, was compared.

Publishing your dataset

Publishing a dataset means giving it addresses that will still work in twenty years, a website where people and software can find each place, and a record in a repository that gives it a DOI. PLATO tools prepares all of that from the dataset itself, in four parts. On the page, choose the part under to publish it and press Prepare; the release’s name, the previous release, the GitHub repository and the rest are under Options. From the command line, each part is publish followed by its name.

Nothing leaves your computer. Neither the page nor the command line sends anything anywhere: they write files, and uploading them, depositing them or opening a pull request is for you to do. The page gives each result as a file (a folder comes as a zip); the command line writes folders.

Every part checks the dataset first, and writes nothing to publish from a dataset that has problems. It needs to know the dataset’s base address: base_uri in the about sheet, or uriSpace in PLATO JSON. Record it there rather than giving it for each run (on the page, “Web address for your identifiers”; on the command line, --base), so that every part, and every later release, uses the same one.

The four parts, in order

  1. Report says what the dataset’s description still lacks to be FAIR: a title, a description long enough for search engines, authors with ORCIDs (whose check digits are tested), a licence given as its web address, a version, what the dataset covers, and a base address that will last. It counts the checks passed, and writes the deposit files (below). Put right what it lists, in the about sheet or the gazetteer header, and run it again.

  2. Mint writes a copy of the dataset in which every attestation has a permanent address of its own, as PLATO JSON Lines (its name ends -with-ids.jsonl). This copy is what you publish, and what the other parts read.

  3. Site makes a website for GitHub Pages from that copy, with the workflow that publishes it.

  4. w3id writes the redirect rules that send your w3id.org addresses to the site. Only for a published dataset whose base is a w3id.org address.

npx github:pelagios/plato-tools publish report my-tables/
npx github:pelagios/plato-tools publish mint my-tables/ --previous my-gazetteer-1.0.jsonl
npx github:pelagios/plato-tools publish site my-tables-with-ids.jsonl --repo my-project/my-gazetteer
npx github:pelagios/plato-tools publish w3id my-tables-with-ids.jsonl --repo my-project/my-gazetteer --maintainer my-github-name

A part that finds problems says so and ends with exit status 1, like Check.

Addresses

Every address is made from the base, which always ends in /:

What

Address

The dataset, and its home page

<base>

A place

<base>place/<id>

A source

<base>source/<id>

An attestation

<base>place/<id>#a-<hash>

A frozen release

<base>release/<name>

The latest dataset, to download

<base>download/<file>

The addresses of places and sources are PLATO’s own rule, set out under spreadsheets to RDF and in step 7 of a first dataset; the rest are how PLATO tools lays out what it publishes. A release is named with --release (on the page, “Release name”); its dataset’s own address (@id) should then be the release’s, with isVersionOf the base, and the report says what to set.

An attestation’s address is part of its place’s: #a- and the first eight digits of a hash of what the attestation says. So minting the same data again gives the same addresses, even from spreadsheet tables, which have no column for them. An address an attestation already has is never changed. Give the previous release with --previous (on the page, “Previous release”), and every attestation it published keeps the address it had there.

Keep the identifiers of places and sources to letters, digits and - . _ ~, in one part (no /), not starting with ., not ending in .jsonld, .ttl or .html (w3id reads those endings as a request for that format), and never two that differ only in capital letters: a website cannot serve any other as a file. While the dataset is a draft the report counts any other as a problem; once it is published its addresses are frozen, and the site lists such places as held only in the downloads.

What “published” commits you to

A dataset is published when its status is published. From then on:

  • its attestations are only ever added to (see comparing two versions);

  • the addresses of its places, sources and attestations are frozen;

  • minting with --previous checks the new version against the previous release, as Compare with the earlier version does, and writes nothing if anything published was deleted or changed. Against a previous release that was still a draft, it writes the copy and warns you what would be refused once that release is published.

Until then, the site says on every page that the dataset is a draft and not to be cited, and asks search engines to leave it out.

The deposit files

Report writes a folder (its name ends -deposit) of files made from the dataset’s description, for a repository that gives it a DOI. Its README.txt says what to check and fill in before you deposit.

File

What it is for

Where it goes

.zenodo.json

Zenodo’s description of the deposit

At the top of the GitHub repository the dataset is released from, where Zenodo reads it with each release

CITATION.cff

How to cite the dataset

At the top of the same repository, where GitHub shows Cite this repository from it

datacite.json

DataCite’s description, for a repository that registers DOIs itself

Sent to that repository, with the DOI added

Once Zenodo has given the dataset a DOI for all its versions (the concept DOI), run the report again with it (--concept-doi, or “Concept DOI” on the page) to put it in CITATION.cff and datacite.json, and give it to site too, which shows it on the home page.

The site

Site makes a website with a page and a JSON-LD document (.jsonld) for every place and source, at the paths the addresses above lead to, and optionally Turtle as well (--turtle). Each attestation with an address is marked on its place’s page, so its address opens the page at it. The home page describes the dataset in the form search engines such as Google Dataset Search read, and offers the whole dataset to download as PLATO JSON Lines, N-Triples and spreadsheet tables.

Beside the site, a second folder (its name ends -repo) holds what goes into your GitHub repository: a GitHub Actions workflow, .github/workflows/pages.yml, and a README-agora.md saying how to set it up. Once the workflow is committed and the repository’s Pages settings have Source set to GitHub Actions, every push that changes the dataset rebuilds the site and publishes it. The workflow runs the same version of PLATO tools that made it, so the site is made the same way every time, and publishes nothing if the dataset has problems. The site itself is never committed.

So what you commit is the copy that mint wrote, with its attestation addresses. The workflow never makes addresses: made there, they would be made again on every run, and an address that changes is no address at all. For a published dataset, the workflow stops if any attestation has no address. Before you push, look at the site on your own computer: the command line writes it to a folder (its name ends -site), which you can open through a local web server, such as npx serve my-tables-with-ids-site.

If the dataset is not at the top of your repository, say where it is with --dataset-path. The page cannot tell which folder spreadsheet tables were chosen from, so for tables it guesses, and says so: correct the path in the workflow, or make the site from the command line.

Size. GitHub Pages serves at most 1 GB for a site, and gives up on a deployment that takes more than ten minutes. PLATO tools estimates the site’s size before it writes anything. Past the limit the page stops and says what to do; the command line writes the site anyway, with a warning, for hosting elsewhere. To stay within it, leave out Turtle, or publish a subset of the places with --only, a file listing their identifiers, one to a line:

npx github:pelagios/plato-tools publish site my-tables-with-ids.jsonl --repo my-project/my-gazetteer --only places-to-show.txt

The addresses of the places left out lead to the site’s “not found” page, which points to the downloads, where every place is.

How long your addresses last

That depends on the base, and the report grades it:

  • A w3id.org address (https://w3id.org/my-gazetteer/) passes. It is a permanent redirect: if the site ever moves, the rules are changed and every address still works.

  • A domain of your own (https://gazetteer.example.ac.uk/) is a warning: the addresses last as long as you keep the domain and its site. When the base is at the root of the domain, the site carries the CNAME file GitHub Pages needs; set the same domain in the repository’s Pages settings and point the domain at GitHub as GitHub’s documentation says. GitHub Pages serves a custom domain only from the root of a site, so a base further down the domain gets no CNAME, and the report says why.

  • A GitHub Pages address (https://my-project.github.io/my-gazetteer/), a local one or a stand-in such as example.org is a warning while the dataset is a draft and a problem once it is published. A github.io address changes if the repository is renamed or moves to another owner, and every citation of it breaks. Use it to try things out; publish under a w3id.org address instead, with the site still on github.io behind it.

Registering a w3id.org address

w3id.org gives permanent addresses by redirecting them, under rules kept in a public GitHub repository, perma-id/w3id.org. A new name is added by a pull request there. w3id writes everything that pull request needs, into a folder whose name starts w3id-, and only for a dataset that is published and whose base is a w3id.org address. It needs the GitHub names of the people who will look after the name (--maintainer, once each; “w3id maintainers” on the page), and where the site is: the repository (--repo), or the site’s address if it is somewhere else (--site-url).

The folder holds the rules and a README for w3id.org, the pull request’s title and text (PULL_REQUEST.md), a list of real addresses of the dataset with what each should answer, and a script, test-w3id.sh, that asks for each. STEPS.md goes through it in order:

  1. Test the rules on your own computer before anything else, in a local copy of the web server w3id.org runs (with Docker), with sh test-w3id.sh http://localhost:8080. Every line should say PASS.

  2. Open the pull request yourself, from your own GitHub account, with the test results pasted into its text. PLATO tools never opens it: once it is merged, the addresses are public and meant to be cited for good.

  3. Once it is merged, test the live addresses: sh test-w3id.sh.

The rules send a browser to a place’s page and any other client to its JSON-LD, and a request for .jsonld, .ttl or .html to that file. The site must be live before the rules are merged, since they only send people there.