How would I process 420,000+ documents by splitting them up into 10,000 document batches?

Viewed 32

We need to add a couple roles to the permissions of each document in a collection. Below is the current code (changed for pii). How would we change the code to process 10,000 record batches? And would that cut down on the space needed? Or is there a quicker way to change permissions on the documents?

declareUpdate();
let aUri = '';
const uris = cts.uris('',[],cts.collectionQuery('DataCollection'));
for (aUri of uris) {
   xdmp.documentSetPermissions(
     aUri, [
            xdmp.permission("data_user_a", "read"),
            xdmp.permission("data_user_b", "read")
           ]
   );
};

We are currently getting an error:

Expanded tree cache full on host marklogic-hostname uri /company/data/document_4534543.json

There is 75 gig free on the drive.

1 Answers

Trying to process that many documents in one transaction will blow out the expanded tree cache, as you have found.

The easiest way to chunk out those updates in batches would be to spawn the work as separate transactions that get executed on the task server. With XQuery, you can use xdmp:spawn-function() with an anonymous function.

Something like this (I'd use a smaller batch size than 10k):

let $page-size := 1000
let $uris := cts:uris('', (), cts:collection-query('DataCollection'))
let $pages := (count($uris) idiv $page-size) + 1
return 
  for $page in (1 to $pages)
  let $start := (($page - 1) * $page-size) + 1 
  let $end := $page * $page-size
  let $uris := subsequence($uris, $start, $end)
  return
    xdmp:spawn-function(function(){
      for $uri in $uris
      return
        xdmp:document-set-permissions($uri, (
          xdmp:permission("data_user_a", "read"),
          xdmp:permission("data_user_b", "read")
        )
      )
    })

JavaScript modules don't have an xdmp.spawnFunction() equivalent. Though, you could invoke xdmp.spawn() and execute an installed module.

However, there are some disadvantages of trying to spawn out giant batches of work on the task server. You can only move as fast as that one task server processing one transaction at a time, and you still run the risk of blowing the Expanded Tree Cache if you happen to use a batch that is too big or hit a pocket of really large documents.

Big batch jobs are better run using batch tools, such as CoRB. It was designed for this purpose. You can adjust the THREAD-COUNT and BATCH-SIZE options, and can execute your JavaScript modules without having to install them in a modules database.

A sample properties file for a job like this would look like:

THREAD-COUNT=32
URIS-MODULE=INLINE-JAVASCRIPT|const uris = cts.uris('',[],cts.collectionQuery('DataCollection')); fn.insertBefore(uris, 0, fn.count(uris));
PROCESS-MODULE=INLINE-JAVASCRIPT|declareUpdate(); var URI; xdmp.documentSetPermissions(URI, [xdmp.permission("data_user_a", "read"), xdmp.permission("data_user_b", "read")]);

I showed putting the code inline in the options file, but you could tell it to read from a file using the filename and |ADHOC suffix, or the path to a module in the modules DB.

and then would execute the job with java command, gradle task, etc.

java -cp "lib/*" -DOPTIONS-FILE=my.properties com.marklogic.developer.corb.Manager xcc://username:password@localhost:8000/database 
Related