What actually happens when a spreadsheet becomes a catalog
The messy middle of migrating a library off Excel — what imports cleanly, what does not, and why a partial catalog is worth more than a perfect one you have not finished.
· The Bookleaf team
Every library we have watched move onto Bookleaf started from a spreadsheet, and every one of them believed their spreadsheet was worse than it turned out to be.
It is worth describing what actually happens, because the anticipation is usually the hardest part.
The file is fine
Librarians apologise for their spreadsheets. They are almost always usable.
The columns are named idiosyncratically — Pamagat, Book Title, TITLE OF BOOK — and that is irrelevant, because you map them yourself during the import. The order does not matter. Extra columns do not matter; leave them unmapped.
What matters is much narrower than people expect: a header row at the top, one row per title, and no merged cells in the data.
Merged cells are the actual enemy
If an import goes wrong, this is why. A merged cell shifts every value after it sideways, so authors land in the publisher field and years land in the subject field. The rows import successfully — they are just wrong, which is worse than failing.
Merged cells usually appear where somebody grouped a section: one merged cell reading “FILIPINIANA” spanning the rows beneath it. Unmerge, put the value in every row or delete the column, and re-import.
The three columns that need a moment
Copy counts. 3 imports. 3 copies, three, and 2 (1 missing) do not. The parenthetical is
real information, which is exactly why it ends up in the quantity column — but it needs to move
somewhere else.
Years. 1998 imports. c1998, n.d., and 1998-2001 do not. This one is worth fixing rather
than dropping, since publication year does real work in a catalog.
Accession numbers. These identify one physical copy, so duplicates fail. When they do, it almost always means the same book is listed twice in the spreadsheet — which is a discovery about your data, not about the software.
The part nobody expects: authors
The import links authors to authority records, so headings stay consistent. But it can only work with what the file contains.
A spreadsheet maintained by several people over several years will contain Rizal, José, Jose Rizal, and J. Rizal, and you will get three authority records. Nothing is broken — merge them — but it is the first moment where a spreadsheet’s tolerance for inconsistency becomes visible as a concrete pile of duplicates.
It is also the clearest illustration of what a catalog does that a spreadsheet does not. A spreadsheet will happily hold three spellings forever. A catalog makes you decide, and then keeps the decision.
Import it before it is ready
The strongest advice we have, and the hardest to follow: import the messy file now.
The alternative — clean the spreadsheet first, then import — sounds responsible and reliably fails. Cleaning a few thousand rows in Excel takes weeks, morale runs out around row 400, and the library runs on paper for all of it.
Importing first gets you a working catalog in an afternoon. It is imperfect. But you can circulate from it immediately, and every book that passes the desk is a chance to fix its record with the physical object in your hand — which is faster and far more accurate than fixing it in a spreadsheet from memory.
A partial catalog is not a failed catalog. It is a catalog that started.
The import documentation has the specifics, and the first week guide has the order to do things in.