A literature is a moving target. Run the same search a few months
apart and the result will have grown, and perhaps lost a record that was
re-indexed. This article shows how to see exactly what changed and how
to merge retrievals safely. It runs offline on the bundled
example_records, a corpus of 138 real journal articles the
package ships because ‘Scopus’ records may not be redistributed. That
corpus is a complete harvest of one query from 2015 to 2024, so a pull
that stopped at 2023 and a later one that reaches 2024 are both genuine
slices of the same search.
The first retrieval ran at the end of 2023 and returned everything published up to then.
[1] 124
A year on, the search is repeated. It now picks up the 2024 papers, and one record that was present the first time has since been re-indexed and no longer matches.
[1] 137
scopus_diff_dois() reports which DOIs were added,
removed or unchanged between the two retrievals, and prints the counts
in each category.
<scopus_doi_diff> 14 added, 1 removed, 112 unchanged
# A tibble: 127 × 2
doi status
<chr> <fct>
1 10.1002/adfm.202315137 added
2 10.1002/asia.202400548 added
3 10.1002/slct.202302535 added
4 10.1016/j.cej.2024.148822 added
5 10.1016/j.diamond.2024.110842 added
6 10.1016/j.isci.2024.111696 added
7 10.1016/j.jallcom.2024.175000 added
8 10.1016/j.jallcom.2024.177248 added
9 10.1016/j.jpowsour.2024.234127 added
10 10.1016/j.jpowsour.2024.236149 added
# ℹ 117 more rows
The newly indexed papers come back as added, the records
present both times as unchanged, and anything dropped from
the later pull as removed. The counts work out at fourteen
added, one removed and 112 unchanged. Fourteen are added because that is
how many of the 2024 papers carry a DOI, and the re-indexed record takes
the unchanged count from 113 down to 112. Records without a DOI cannot
be tracked this way at all, which is one reason to prefer the ‘Scopus’
identifier when there is one.
To act on one category, filter the table, which is an ordinary tibble.
| doi | status |
|---|---|
| 10.1002/adfm.202315137 | added |
| 10.1002/asia.202400548 | added |
| 10.1002/slct.202302535 | added |
| 10.1016/j.cej.2024.148822 | added |
| 10.1016/j.diamond.2024.110842 | added |
| 10.1016/j.isci.2024.111696 | added |
To keep a cumulative set across retrievals, combine them.
scopus_combine() renumbers the records and, with
dedupe = TRUE, keeps each one once by ‘Scopus’ identifier
or DOI, so the records the two pulls share are not doubled.
[1] 149
That is 149 rows for 138 distinct articles, and the gap is instructive. These records carry no ‘Scopus’ identifier, never having come from ‘Scopus’, so de-duplication falls back to the DOI. The eleven that arrived without one have no key to match on, and so survive in both copies. A live harvest carries an identifier on every record, so the same call on two real pulls returns each article once.
The base c() method concatenates record sets directly,
renumbering but without de-duplicating, so it is the building block that
scopus_combine() adds the duplicate handling to.
[1] 261
Saving each retrieval lets you compare against it next time. The
.rds form round-trips exactly.
path <- file.path(tempdir(), "baseline.rds")
write_scopus_records(baseline, path)
identical(read_scopus_records(path), baseline)[1] TRUE
A live retrieval also carries the date it was taken, as the
retrieved_at attribute, together with the
scopusflow version that took it. That matters for this
workflow in particular, because citations is a snapshot
value that keeps moving, so a difference between two pulls only means
something once you know how far apart they were. Both survive the
.rds form and neither survives .csv, which is
a table of columns. The bundled corpus used here carries neither, which
is why it round-trips identically above.
In a live setting the later retrieval would come from the API, where here it is a slice of the bundled corpus, and everything else would be as above. Both pulls could then say when they were taken.