Skip to content

Chapter XI: Strait of Magellan.--Climate of the Southern Coasts (4)

Text size

All of the above at different times. I am not an avid television watcher and would rather do some "work" (or should I say "pleasure") for PG much of the time.

HOW MANY DIFFERENT KINDS OF WORK, OR DIFFERENT BOOKS, HAVE YOU DONE?

Because I have converted many books from work already on the internet, I have covered quite a range, though I haven't actually scanned and proofed too many books. Those that I have done have been Australian historical works. But I have rounded up books on philosophy, aboriginal legends, and several novels. Since many internet sites come and go, I am interested in "grabbing" etexts and posting them at PG in case the site disappears from the internet. It has become a pastime in itself. I recently discovered "South Wind" by Norman Douglas, a book which caused quite a sensation when it was first published because it portrayed a bohemian lifestyle. Ironically, I used to have the book in my home library, but dispensed with it when I needed space. Now it is at PG and I can get it whenever I want it.

WHAT DO YOU LIKE ABOUT THE PG PROCESS?

The democratic, helpful, friendly approach of all the people involved is one of the things I like best. I have "met" so many wonderful people, without having to "live" with them, if you know what I mean. Not long after I started associating with PG, Michael Hart posted an e-mail to the volunteer discussion group, advising of the death of a long-time volunteer. It seemed like she had been one of the "family".

One really needs to be indifferent to praise and the prospect of reward to start volunteering for PG. There is certainly no money in it. However, one quickly finds that there is a community of people out there with a common interest, and with the same outlook and the same interest in doing a job well, without tangible reward. There is no lack of praise though, and one soon finds that one is not indifferent to it.

WHAT DO YOU DISLIKE ABOUT THE PG PROCESS?

There isn't much that I don't like. Nothing worth mentioning, anyway.

IS THERE ANYTHING YOU'D LIKE TO SEE PG DOING DIFFERENTLY?

There are a few things, however since I don't know all the reasons for some things being done the way they are, and because everything is done by volunteers anyway, I wouldn't like to canvass them here. To have produced nearly 5,000 etexts over more than 30 years is testament to the fact that most things are being done "right".

IF ONE OF YOUR FRIENDS APPROACHED YOU TO ASK ADVICE ABOUT HOW TO GET STARTED CONTRIBUTING TO PG, WHAT WOULD YOU TELL THEM?

I would spend some time with him/her and work through some of the issues. I know that I would have benefited from that approach. I would gradually introduce her(him) to the different issues which need to be addressed and find out exactly what her expectations were, and try to help her in fulfilling them.

WHAT WOULD YOU EXPECT PG TO BE LIKE IN FIVE YEARS? TEN YEARS?

Much the same as it is now, I hope. After all, the goal will continue to be to provide "fine literature digitally re-published". Though I expect that, like other organisations, it will continue to evolve in response to new challenges and opportunities. Ten years ago, who would have thought that there would be 5,000 etexts posted; that there would be volunteers operating an online proofreading site; and that there would be a volunteer writing free software to read PG etexts? The rapid growth of PG over the last few years will present many challenges for the future.

Writing of etext readers, I am reminded that I recently joked to a volunteer that I wanted him to write software for reading etexts, whereby a hologram would appear on the inside of my eyelids so that I could read etexts with my eyes closed. Who knows, it might be possible. However, whatever advances in technology occur over the next ten years, one thing is certain: the work of all the volunteers to date will ensure that there is an amazing library of ebooks available covering creative works by some of the greatest minds who have ever lived. Future readers of PG ebooks will have been given a wonderful gift by the many volunteers who have contributed to PG over the decades.

Project Gutenberg of Australia

On the wall in a colleague's office was pinned a piece of paper on which was written a quotation. I don't recall now what it was and the colleague has been gone for some time and has taken the paper with him. However under the quotation the author was acknowledged as "Prince Machiavelli". I had a vague idea that the quote actually came from "The Prince" by Nicolo Machiavelli, and wondered how I could satisfy my curiosity. Then I remembered reading about Project Gutenberg and decided to see if the book was posted on the PG site, though I didn't really expect that it would be. Needless to say, the etext WAS there and I was able to download it and read it in its entirety, due to the time spent by John Bickers and Bonnie Sala (their names appear at the beginning of the etext) in preparing it for PG. Interestingly, there were other works by Machiavelli there, which I hope to get back to one day.

Later, when I e-mailed PG and expressed an interest in volunteering I was, because I said that I was Australian, referred to Sue Asscher, the Australian Production Director for PG. Sue asked me to proofread "A Vindication of the Rights of Women" by Mary Wollstonecraft. Also, about this time, a journalist had contacted Sue with regard to a story being prepared for PG. He wanted to contact some volunteers to ask why they were interested in PG. Sue referred the journalist to me, with my permission of course, and one of his first questions was "Is there much Australian content on PG?" After I had checked the PG etext list I could only reply "not much".

So I decided to start creating etexts by Australian authors, for PG. Sue Asscher pointed out that there were many eligible Australian works already in the public domain as etexts, so I started rounding up etexts and matching them with books which had been published before 1923, so that they could be posted at PG. Then I started creating etexts myself, for works I could not find already on the internet. My sister had given me, many years ago, a book by Geoffrey Dutton titled "Australia's Greatest Books", so I decided to start working my way through the eligible titles from the list of about one hundred books reviewed by Dutton. I had already found a number of them on the internet and some were already at PG. But there were still a "few" to be done. There still ARE a few to be done, if anyone is interested in helping.

Then Sue Asscher again had a hand in setting the direction I would take by asking me to proof an etext of "Animal Farm" by George Orwell, whose work had recently entered the public domain in Australia. We didn't know where we would post it, as it is not in the public domain in the US, but I agreed to proof it as I had read it many years ago and enjoyed it.

About this time, I also decided to make up a personal web site. Being a software developer, people were always asking me about the internet and web sites, in the mistaken belief that I knew ALL about computers. I decided to get an idea of how web page design and web site management worked by creating a site that listed all of the "Australian" content at PG. When I couldn't find anywhere to put the Orwell, which I had recently proofed, I decided to create a page on my site for etexts in the public domain in Australia, so that Australians and internet users in other countries with similar copyright laws, could read and/or download them.

Michael Hart, the founder of PG, was quick to interest me in creating an "official" PG site in Australia. After registering a business name, getting a domain name and finding a sponsor to host the site, Project Gutenberg of Australia was up and running.

It all happened very quickly, and as with many things which happen in one's life, it all seems to have come about by serendipity. Even the site's motto "A treasure-trove of literature" was stumbled upon by chance when I looked up, in connection with another unrelated matter, the word "treasure-trove" in a dictionary, to ascertain if the word was hyphenated. Imagine my surprise to find treasure-trove defined as "treasure found hidden with no evidence of ownership". That EXACTLY defined the literature found on PG.

My own association with PG resulted from the culmination of a life-long interest in books and literature and an equally strong interest in computers. Every volunteer brings his/her own particular interests and skills to PG and that, together with the democratic approach taken by the small executive team, is what makes PG the strong, co-operative organisation that it is. My interests and skills, and a generous dose of serendipity, led to the creation of Project Gutenberg of Australia.

Dagny

I discovered Project Gutenberg in 1996 and immediately wanted to help because I love books and wanted everyone to have access to all the wonderful books that, even today with Internet searching, are difficult to find or very expensive when you do locate them.

I began by proofing a few works but what I really wanted to do was share my Balzac collection with other fans. I discovered Balzac in the 1970s and recall my frustrations in trying to find more than a dozen stories of the over one hundred Balzac wrote. It was over a decade before my husband discovered a complete set at a used bookstore while on vacation. Unfortunately, not everyone is so lucky.

With the first few stories I typed for Project Gutenberg I worried about everything: should I correct a type-setting error, leave it, footnote it, etc. This took a long time and involved a lot of correspondence. Now, my idea is to make the text as readable as possible. For me that means correcting type-setting errors I notice. Others prefer to leave them intact. In the end, I don't believe the readers care. I have found them generally to be very grateful to have found some treasure they had been seeking. In some cases of an author's more obscure works, they didn't even know the book existed, a rare find indeed for them.

It is so satisfying to receive an e-mail from someone thanking you for all your hard work. Most readers don't take the time to write but true fans often do and they make it all worthwhile. I have even met people in this way that went on to become a Project Gutenberg volunteer themselves because they wanted to give something back to the Project from which they had received so many pleasurable hours.

Gardner Buchanan

SOURCE MATERIAL

First of all, there is the issue of what texts I choose to do. For me, this is fairly simple. I'm a bit of a small-time book collector already, and have a personal theme: "Canadian English Literature" and "Canadian English-Language History". I have no trouble whatsoever in coming up with submissible editions of works that fit this theme somehow. Nevertheless there are specific authors and works that I'm not having luck with, so I'm still making the rounds of the used book shops regularly and picking up all sorts of stuff.

Eligible volumes have typically cost me $10.00-$150.00 for a collectable edition, or $0.50-$15.00 for a recent paperback edition or garage-sale item. I paid $0.50 for a eligible, but not very collectible copy of Glengary School Days by Ralph Connor at a garage sale. As it turns out someone has beaten me to it--it has been in the collection since 2001. Sometimes if I'm contemplating picking up a more expensive book that I don't already have a personal interest in, I'll go back and double-check The Online Books page to see if someone has already submitted the book.

Another way I obtain texts is from the Early Canadiana Online archive. They host page images of quite a large collection of old books written in or about Canada, or written by Canadians. The page images are reasonably well suited to OCR.

I tend to produce E-texts two different ways. One way is to submit page images to Charles Franks who runs Distributed Proofers and let him worry about bulk-OCR'ing. I then manage the distributed proofing, which is a fairly low-effort business. The other way is to scan, OCR and proof all by myself. I'm currently averaging two of my own projects to every Distributed Proofer one.

SCANNING AND OCR

I have an very slow parallel-port scanner, a UMAX Astra 2000P. It sucks mightily. I'd rate it a 2 out of 5, if it wasn't acting up--creating a black bar across the page, part way along--so I have to scan books a certain way around to avoid having the bar land in the text. As it sits now, it's in 0.5-1 territory. It is glacially slow at the best of times, and due to being a parallel port model, locks up my whole computer during the scan.

Nevertheless, it is completely adequate to my needs for PG work. I've scanned more than a dozen books on it, and it's done yeoman service--despite its warts. Scanners like this one can be picked up used for $30, and are worth the money.

The way I work when I'm producing a book myself, is scanning and proofing page by page. I do the scans two-pages-up, then OCR, proof and copy the pages to a working document, before going on to scan the next pair of pages.

My scanner came with two OCR "packages": Omnipage something-or-other which I was never able to install, and Recognita Standard 3.2.7. I use Recognita, and for 300dpi scans I do, it is adequately fast and accurate. It is a no-frills package, and DOES make many mistakes, but it is entirely useable for my purposes. I rate it 2 of 5.

I've used the Abbyy FineReader 5.0 try & buy. This is a magnificent OCR system. It handles huge batches and is fast and astoundingly accurate. I rate it 5 out of 5. Unfortunately it costs about $million to patriate a web-bought item into Canada, and while priced at a very reasonable US$100.00, would cost me about CAN$600 after exchange-rate, brokerage fees, shipping, more fees, taxes, service charges and more taxes (on the fees).

I could buy Omnipage off-the-shelf here, but frankly if I can't get Abbyy, I'll stick with Recognita.

As I scan each page, I paste it into Windows-95 Wordpad. Sometimes I also do some proofing in Wordpad, but mainly I proof, fix quotes, M-dashes and paragraph breaks in the OCR program before copying to Wordpad. I like to keep the page boundaries intact, and I mark them in my Wordpad document like this:

:
:
kjdk ldjd ll;llkj dklj dklj
kjdk ljd llllkj klj dklj

page 354

kjdk ldjd lll;;llkj dklj dklj kjdk ldd lll;;llkj dklj dklj kjdk ldjd ll;llkj dklj dklj kjdk ljd llllkj klj dklj

page 355

kjdk ldd lll;;llkj dklj dklj kjdk ldjd ll;llkj dklj dklj kjdk ldd lll;;llkj dklj dklj kjdk ljd llllkj klj dklj : :

At this point I also fix-up hyphenated words that straddle page-boundaries. I note paragraphs that start in a new page and mark them with <p>, and I note indented or block-quoted sections and mark these with <in>..</in>. This helps when I go back to format it since I can easily see where the special cases are.

Wordpad handles large documents reasonably well and will grok UNIX files (ie: <LF> only, not <CR><LF>). For this it rates 3.

PROOFING AND FORMATTING

When the whole text is assembled, whether by myself or by Distributed Proofers, I use about the same process for formatting and final proofing.

I use MS-Word 95 to do a spellcheck. This I rate 3 out of 5. I do a select-all, and language appropriately - for me, usually UK rather than American English. I wish I had a Canadian English dictionary for Word 95, but have not needed one badly enough to actually look. Word has a pretty good spell checker and the custom dictionaries are easy to muck around with. I use a custom dictionary for any big project - I have one for Chronicles of Canada, and different one for all the John Richardson books I've done.

At this point in my personal process, I abandon Windows and go over to FreeBSD.

I use vi (rated 9 out of 5) to do a number of hacks. I search for and fix up hyphenations that were broken (peer- less) and such like. I also search for and fix some OCR special case errors like 'you'->'yon' and 'be'->'he'. This latter sometimes requires a while, just to step through all the be and he's to see if they're right.

Still in vi, I next use some incantations to run the UNIX 'fmt' command on each paragraph to get it reformatted. I use:

fmt -55 60

Fmt gets a 3 out-of 5 for what I need it for. It double spaces after sentences, which--although it is probably the right thing to do--is not the PG convention (for me at least). It also adds a space when joining lines with an M-dash. I go back and fix both of these using vi. I take into account the <in></in> tags and manually format accordingly at this point.

As I reformat, I give the text it's final proofing. I'll have the original text in-hand at this point, and will use the page markers (remember them) to figure out where I am. As I reformat, I delete the page markers and other markup. When I'm finished this step, the book is almost done.

Next, I use Gutcheck 0.2 (5 of 5, for intended purpose - way to go Jim!) to check for all the things it checks for. At this point I usually get something like 50 hits, of which 30 are real. I'm then back in vi, and fix up all those problems. Finally, I'm done.

As I go along, I tend to keep various versions of the document. I'm at version 27 of 'The Imperialist' right now. Each scanning editing, spell checking or whatever type of session gets a new version: imperialist_12.txt, imperialist_13.txt,... At various times I might find it useful to use 'wc', 'grep' and 'diff' to figure out what is going on, where a word appears or whether I deleted something I didn't mean to.

HARVESTING PAGE IMAGES

I mentioned above that I sometimes work from page images that I obtain from the web. There are several archives around that hold eligible materials as page images that you can easily download and OCR. I personally have worked mainly with the Early Canadiana Online archive.

After a bit of poking around with the web interface to this collection, I have been able to work out how the individual pages are numbered and organized. I have written some shell scripts that I can use to fetch all the pages of a volume and convert them from GIF to TIFF format. Harvesting a 200 page book takes a few hours.

Once I have all the pages, I have to do some work with an image editor to get them ready for OCR. I use Corel PhotoPaint 7 to crop each image to just the text area and to remove the black bands at the sides due to the spine or whatever. The page images are often made from microfiche, and dust marks are common as well. These I can sometimes edit out with PhotoPaint.

Because some of the page images, or certain sections thereof, can be completely unreadable, I often find myself either tracking down a modern edition or visiting a local university library to find a copy of the book to look up a few paragraphs or passages that are not readable in the images. Even having to do this, I find that the capture of images from the archive is still a big time saver, and allows me access to an edition that would otherwise be totally inaccessible.

Having gathered the images and prepared them for OCR, I next submit them to Charles at Distributed Proofers, or handle them myself, using the same process as if I were scanning them.

DISTRIBUTED PROOFERS

I've done several books using Charles Franks' most excellent Distributed Proofers web application. I tend to choose DP when I don't have the personal time to read and proof a volume myself, or when the poor quality of the text defies the ability of my (not very good) OCR package.

When scanning for DP, I still scan images two-up. I then have a collection of shell scripts that cut the page images in half to produce single-page TIFF files. I then use a manual procedure with Corel PhotoPaint 7 - if required - to fix up skewed pages or ones with black margins. For the most part, page images that I scan myself are registered exactly enough in my scan area that the page images don't need to be edited.

Page images that I've harvested from a web archive do have to be fixed up before they can be used by DP.

Charles, I believe, prefers that as a project manager I would deal with my own OCR. He has, however, been kind enough to run several batches of page images through his OCR setup for me to good effect. I believe he uses Abbyy Finereader, and my procedure for submitting pages to Charles is to run a subset of the pages I intent to send him through a demo copy of Finereader to make sure that the results are vaguely acceptable. If everything looks good, off it goes.

When the project has run its course with DP, I download the completed text and proceed to format and re-proof it, for the most part, as if I'd scanned and OCR'd it myself.

Jim Tinsley

How I (eventually) got started.

Five years ago, I was the most clueless newbie ever to try volunteering for PG. If you're feeling lost about how to help PG, you can be sure that you're not alone! And if I can write PG's first complete FAQ after my bad start, you can surely do better! :-)

Back in 1997, the web site existed, but there were no FAQs, no Volunteers' Board, no gutvol-d, no Distributed Proofing sites. I started by making a donation and e-mailing Michael, suggesting that I could help out with small jobs, or programming. I didn't get any, and I had no idea what, if anything, I could usefully do by myself.

I looked up the in-progress list at the time, and e-mailed a few people who were listed as working on books, offering to help. None of them were still working on the books. (We no longer show people's e-mail addresses on the InProg list.) I still had no idea how to get eligible books, no scanner, and no idea how to approach producing an etext.

I subscribed to the monthly Newsletter, and just read it for a year. In a "Project Gutenberg Needs YOU" edition, Dianne Bean, the U.S. Director of Production at the time, was given as a contact. I e-mailed her, and finally things started happening.

She sent me a short piece to second-proof, and explained that I should just fix whatever needed fixing. I returned it, and she introduced me to Bill Brewer, who was, at the time, scanning Wisters like they were going out of style. He and I formed a scanning/proofing team for a while.

How I began producing, and my problems with scanning and OCR.

I had some ideas for books I wanted to produce, but I couldn't find them locally, so I turned to the Internet, and discovered how easy it is to find and buy used books on-line.

I bought a HP flatbed scanner. It came with freebie OCR software-- "PrecisionScan"--with images and OCR all in the same interface.

I scanned my first book, which fortunately had large, clear text, and the OCR made a reasonable job of it, according to my standards at the time, which were that getting any text at all without typing was a form of magic :-)

I now know that I could have made a better job of it if I had pressed the spine down hard, either closed the top to keep out ambient light or darkened the room, and made each scan a bit more exact. I'm much better at flatbed scanning now.

My PrecisionScan software _did_ recognize two facing pages, and dealt with them correctly, though IIRC it put some garbage characters between the pages that I had to remove by hand.

It did require a lot of editing, though, and recently I've gone back over my original text and found lots of mistakes. Partly because of the scan, partly because of my inexperience.

Throughout the editing, I kept having to make formatting decisions in a vacuum, reinventing wheels and applying rules from a HowTo. Now, having read and formatted and proofed and produced so many texts, I just _know_ how to format a text without thinking, and just reading or even skimming a few texts before producing my own would have given me a lot of background and saved a lot of time. I had proofed several books, but never thought to look closely at formatting decisions.

That text took me a month of working most evenings, and a lot of sticktoitiveness. I can really appreciate the effort that a volunteer has to put in to produce their first text by casting my mind back to that month. I think it's the not-quite-knowing-what-you're-doing that's the worst part. I remember being soooo relieved when I sent it off for second proofing.

The guy who took it for second proofing didn't get back to me for a month, and then said that he wasn't going to do it. This was disappointing. I sent it to another guy for proofing. He came back after a few weeks asking some questions. I answered them. After a few more weeks, I followed up with another e-mail. No answer. A few weeks after that, I gave up, and just submitted the file for posting.

The next book I produced didn't have such nice, clear, large type, and the scan was what I would today call abysmal. I'd guess that I retyped a quarter of the book. The less said about that one, the better.

My third book just _would not_ OCR sensibly. The print was very small and faint, and the OCR produced gibberish. Even with my low standards, I couldn't kid myself that this was working. I tried 400dpi, 600dpi. No dice. I might get 10 complete words on a page.

It was at this point that I bought TextBridge. I really had no idea about the difference between the freebie OCR programs they give away with scanners and a genuine commercial product, but I was trying in desperation to get _something_ different that would read this image.

Textbridge was an eye-opener for me. It still didn't make a good job of the bad images, but it made a decent shot at maybe half of them, and having bought it, I tried it on the two books I had worked so hard at before--it gave hugely improved results. The book that had only been about 75% OCRed became 100%, but with some errors. I cursed the time I had wasted making up for the deficiencies of my freebie package.

Since then, I've kept upgrading my TextBridge (I think I started on version 8, now on Millennium) and bought OmniPage and Abbyy as well. I mostly use Abbyy 6 now.

Last time I looked, there were downloadable trials of Abbyy, TextBridge, and OmniPage. Big downloads though.

Last year, I got a new Epson Perfection 1640 scanner to replace my old HP Scanjet. I never had any complaint about the Scanjet itself--it served me well--but the new Epson is faster, has higher resolution, and ADF.

Even better, I now know how to scan. I know how to process 200+ pages an hour while scanning the book flat, two pages at a time. I know how to adjust the settings to scan only the area covered by the book. I try different settings for each new book to see what works.

So much for scanning and OCR. I was a _very_ slow learner in this area.

How I prepare a text now.

I was never quite so bad on the proofing end of things. As an editor, I use Brief in DOS and Crisp (a Brief clone) on Windows. (I mostly use vi on *nix, but I do very little-to-no PG work on *nix apart from an occasional scripting thing that I can do in one line of Perl, but would be annoying on MS).

Now, I'm all for tolerance and equality and respect for the faiths of other people, :-) but I gotta say that for someone who has used a powerful editor, editing with Word or any standard Windows editor is like scratching your nose with a rake.

When I first get the text off the OCR, I have many pages with breaks between them, and usually no line-spacing between paragraphs, but each paragraph indented.

I whip out Crisp, and run a macro to search and destroy all page-breaks and page-numbers and blank lines between, and then another to put line breaks between paragraphs and unindent them. Since I watch this process carefully to avoid messing up quotations, it takes me maybe 15 minutes.

Now I have a basically formatted text. The line-lengths are usually too short, and there are hyphenated words at line-ends that I will need to rejoin, and some that I need _not_ to rejoin. Another macro fixes up the hyphenation. At each hyphen, I just decide whether to rejoin or not. Say 20 minutes, max. Then I rewrap. Another 15 minutes.

So in maybe an hour I have a proofable text, and the really nice part about it is that I've had a flying tour of the text three times, so I've already noticed any peculiarities.

If I've noticed any unusual features like letters or poems that need special treatment, I do it at this point.

To prepare the text for proofing, I just flick through it in Crisp with spellquery on, in US or UK English as needed. This puts a red line under queried words, just as Word does. I spend maybe 5 or 10 seconds per 50-line screenful. I don't expect to catch them all; this is just a quick pass to thin 'em out. I may also catch some formatting issues, but I'm not looking for them.

Now I proofread.

I've tried lots of ways of proofreading. Often it's just sitting at the screen. Sometimes I print out the texts or parts of it, and mark errata with a pen. Occasionally, I get the computer to read the text to me, and I follow along in the book, noting any errors. (This is good when you want very high accuracy - do a replace of ":" with "colon", "," with "comma" and so forth before you start the reader.) Recently, I've tried reading the text on a PDA, and bookmarking the problems.

Whatever way I do it, it takes time. I'm better at it now than I was, but I still tend to miss things like he/be.

Some people swear by particular fonts for proofreading, saying that font X shows "1"/"l" differences more clearly than font Y. I just use Arial or Verdana for printouts and Courier or Fixedsys on screen; the special fonts don't seem to make a difference to me.

So I've finished proofing and made my corrections. Now I leave it sit for a few days. I need to get my mind off it, so that I won't miss the same errors I missed before.

When I come back to it, I'm looking at what software people would call a Release Candidate, and something changes in my head . . . I'm thinking of it in a different mode, not as a work-in-progress, but as a potential finished project. This makes me much more critical, and less willing to accept mistakes.

Usually there are dash-problems to fix up (emdashes as " - " instead of "--") and other minor stuff like that. I do global searches for " -" and "- " and "...".

I do a quick skim though it, sampling paragraphs here and there as a test of its quality. I make any formatting adjustments like chapter line spacing or indenting letters that I might notice.

Then I run gutcheck. Gutcheck is a little program I wrote / write / will-write over the years that complains about common problems in a PG text . . . bad line-lengths, common typos, numbers within words (like the "1" in "wor1d") unbalanced quotations, spaced or unspaced punctuation, non-ASCII characters. I fix the problems that Gutcheck points out.

Again, I switch spellquery on in Crisp, and skim through, more slowly than the first time. This time, I'm looking for _anything_ that shouldn't be in a PG text.

I run gutcheck again, just to be sure.

And off it goes!

The Posting Team

For a couple of years, I churned out a text regularly every two months, spending about 40 hours on each, and took on some occasional proofing, but after I became moderator of the Volunteers' Board, people started referring texts to me for checking or reformatting. This took up more and more of my available PG time, and my own production slowed accordingly.

It was in response to these requests that I wrote gutcheck, which embodies all the standard non-spelling checks I would run on a file. Gutcheck allowed me to spend less time on each text, but still feel reasonably sure that there was nothing glaringly wrong with it.

When Michael formed the Posting Team last year, I volunteered, and it was a natural progression for me, since I was already used to doing a lot of last-minute work on texts.

I found posting to be disorienting and confusing at first; people bombard you with half-scraps of information about books to be posted; some texts need serious work; some texts haven't been cleared, and need to be referred back; some people want special treatment for their texts, which may conflict either with my views or with PG precedents, or both; there are lots of questions. But like every other new job, it just takes time to learn the ropes.

The actual process of posting now takes very little time: I can go through the necessary steps in 3-5 minutes. But posters are the last line of defense against errors, and even the most careful volunteers make them (and yes, we do too!). It takes a minimum of 15 minutes to run standard checks on a perfectly clean file, and it can take several hours to fix up a file that needs help. On average, it takes me about an hour to do my reasonable best for every text submitted.

Apart from posting proper, there are a lot of queries to be answered, many of which I hope I've dealt with in this FAQ, "special cases" that eat as much time as I'm willing to give them, corrections to be made to existing texts, and interminable debates about whether PG should do _this_ or _that_.

Now that the learning curve is past, the problem with posting is that it generates a lot of e-mail and discussion, and eats a lot of time, and is a 7-day-a-week commitment. Having posted over a thousand texts, I'm now particularly interested in ways to improve text quality.

John Mamoun

How to create an e-text efficiently or automatically is an interesting logistical problem. Here is my procedure, which I recently used to make an e-text in about a week, with maybe 6 man-hours of work on my part:

I take the book, and use an x-acto blade to cut out all of the pages. I then feed the pages into an HP 4C scanner with an automatic document feeder accessory attachment that I got from e-bay for $200. I feed it up to 50 pages at a time, and it automatically scans them in.

I work the scanner using software called scan2000, from www.informatik.com (30-day shareware trial period, $50 to register). This program automatically works with the scanner to save each image as a CCITT4 standard format TIFF file. Most importantly, it automatically numbers each page, starting with an initial value you specify (typically 001.tif) and increasing the number of the file name by an increment you specify (typically by 2 pages, since you scan double sided pages; you scan the evens first, then flip the pages over and scan the odds, but you want the page numbers in order, right?). So the scanner outputs, say, 001.tif, 003.tif, 004.tif, etc., then you flip the pages over and re-feed them into the scanner; the even pages are saved as 002.tif, 004.tif, etc., after you tell the program to begin the first of the even page files with 002.tif.

So now I have a bunch of consecutively numbered CCITT4 TIFF files. At this point, I could use a freeware program called cc42 (search for it at www.pdfzone.com) to combine all of the sequentially numbered CCITT4 TIF files into a single PDF file with the pages in order.

Or, if making e-texts, not PDF files, I OCR the pages and save them as corresponding pages like 001.txt, 002.txt, etc. I also use Paint Shop Pro (shareware 30 day trial) to batch-convert the tiff files into GIF file format. I can then upload the GIF files and the correspondingly numbered text files to the Distributed Proofreaders page (http://texts01.archive.org/dp/) to have them rapidly proofread by numerous proofreaders, who finish the task at a rate of 50-100 pages a day per book, very roughly speaking. When done, I then download the text files as a single text file combining all of the files. The upload function on the DP site is tedious, requiring one to upload each file one-by-one, but I spoke to the webmaster recently, and he said there are, with special arrangements, ways to FTP them or even e-mail them to him on CD.

Now, hard returns. It was once a grave problem to fix hard returns so that the text outputted to 65 characters per line. Then I got a freeware program called Clipcase at www.shareware.com. With Clipcase, you select a body of text (about 20 pages or so; any more, and the program crashes) in your word processor, copy the text to the clipboard, then load up Clipcase, paste the text into the Clipcase window, the process the text.

When this happens, all of the hard carriage returns within the text are eliminated, EXCEPT for returns between paragraphs. Then, you select the text, copy it, and paste it into any word processor to process it. I use Microsoft Word. After pasting all of the text into it, I select all of the text, choose Courier New font, 10 point size, and set the margins at 5.5 inches. With this setup, when the text is saved as "Text with layout," the resultant text is 65 characters per line, every line. Setting hard returns is automatic.

Then I spell-check the text, and also skim through it to look for typos and "categories" of errors to tend to occur repeatedly within the text. One common error is having a single dash instead of two dashes, for example:

He lingered-slowly. as opposed to: He lingered--slowly.

Another common error is a space between a period, exclamation mark or other punctuation mark, and the letter that came before it, such as:

Hey ! instead of Hey!

or " Hey, " instead of "Hey,"

I then use the "Find/Replace" command within Microsoft Word to efficiently get rid of these. For example, I might tell it to look for ^w", where ^w means "a white space" and " is a quote. This looks for white spaces before quotes. "^w looks for white spaces after quotes. ^w! means a white space before an exclamation mark. I can also have it look for "any letter"-"any letter," so that it finds single dashes between letters, and then I can decide if I want to replace these with double dashes. By using these kinds of find/replace tricks, it becomes easier to remove typos.

When done, I save as "text with line breaks" and it is done.

That's basically my procedure. 1 week turnaround time and 6 man-hours on my part for a 190k text file...

Ken Reeder

The Story of My Life (as pertains to PG) by Ken Reeder June, 2002

I am currently finishing up my fourth etext, with two more etexts in process, another seven books sitting on the shelf waiting, and a lot of additional books that I would like to do when those are done.

Sixteen months ago I was blissfully unaware of PG and of the world of online books. A couple of things seemed to come together to lead to my involvement with PG. I spent some time helping one of my sons, for a school project, in an unsuccessful search for an online English translation of Pliny's Historia Naturalis. About a year before that I had been tinkering, for no particular reason, with trying to type one of my favorite older sci-fi books into a text file. And I had been thinking, occasionally over the course of a few years, about a series of books to which I was avidly devoted when I was about twelve or fourteen years old, which was widely available then but is relatively scarce now. It was a web search on the name of that author, Joseph Altsheler, which happened to lead me to some couple-year-old messages on the PG volunteers' bulletin board.

I poked around the PG web site a little and thought, hey, I think I could be interested in this. Only a few months before I had, for no particular reason, picked up a clearance-model parallel flatbed scanner (for which I paid $36, including shipping). The scanner package included some OCR software, so I already had the basics needed to scan a book to produce an etext.

So I rummaged around on the PG web site a good bit more, and lurked on the volunteers' board, and figured out that I could find the books that I wanted on Ebay or ABEbooks, and bought a couple of books for $10 or $15 each. I scanned a chapter or two and tried out the OCR, which worked very well. (The OCR software that came with my scanner is TextBridge Pro, which it turns out is one of the more highly-regarded OCR packages, so I was just lucky in that respect because I had no clue. I could see that the OCR software was clearly much better than some DOS software that I had used at work about 15 years ago.)

What appealed to me was that, firstly, it seemed like this was a worthwhile thing to do, with a big plus being that you can do the work from your own home, in your pajamas if you want, in whatever time you can spare. And I thought that, being a detail-oriented software-developer geek kind of guy, that I would kind of enjoy it and also be pretty good at it - actually, I've always had an aptitude for proof-reading.

So I went ahead and mailed in a couple TP&V for copyright clearance, and set out to actually produce my first etext, a 348-page book which I completed in about 10 weeks, start to finish.

For a book with nice clear, good-sized print, I figure that it averages out to about 7 or 8 minutes per page to go through my complete production process. Some of the books that I am working on, with smaller or less-perfect print (and/or other complications) take a little (or a lot) longer.

I feel that I've got my process pretty well set by now. I've put together several little home-made utility programs, written in FoxPro, which assist me. (I've put in some effort to try to adapt some of these for possible use by others, but the problems are that it takes a lot more work to polish software to the point that I feel comfortable letting somebody else pound on it, and the scope of what I think the software ought to do gets bigger every time I work on it, and it's not nearly as enjoyable - for somebody who develops software at work every day - as producing etexts.)

My complete production process, with rough time breakdown, is as follows:

1. Scan the book, 2 pages at a time, about 1 minute per scan (30
seconds per page). (I do not cut the pages out of the book, I
just lay it flat on the scanner and press down on the spine.)

2. Run the BMP file through TextBridge Pro, about 30 seconds per
page. (Again, when working with clear, good-sized print.) I
save the output as text with no line breaks.

3. Run a little FoxPro utility that I wrote that massages and
formats the file a little bit.

4. Do my first-pass proof-read, about 2 minutes per page, combining
the pages into chapters.

5. Run another little FoxPro utility, which checks for some things
that I might have missed during proof-reading.

6. Use MS Word to perform a spelling and grammar check, another 30
to 60 seconds per page.

7. Run another little FoxPro utility (number 3), which inserts line
breaks, then run another one (number 4) which does some more
exception-checking.

8. Do my second-pass proof-read, about 2 minutes per page.

9. Combine the chapters into one big file. Run a couple more little
FoxPro utilities (numbers 5 and 6) which do some final formatting,
checking and analysis.

10. Send the file to Jim Tinsley, who will graciously run it through
his GUTCHECK program which scans for a lot of common errors.

11. Call it an etext and send it in for posting.

My primary goal is to produce a quality etext - I don't particularly care about trying to speed things up. I mean, I don't want to needlessly waste a lot of time, but I look at this as a hobby and I enjoy working on it, so I don't get out my stop watch to see if I can get 20 pages done faster today than yesterday. (When I go out running, then I'm concerned about whether I'm faster today than yesterday.) I generally put in maybe 5 hours a week on PG - actually, it's often easier for me to fit in some PG work on weekday evenings than on the weekend. And it is definitely gratifying when the etext is done and not only does it get posted on PG, but then links and copies pop up in different places like the "Online Books Page", and DMOZ.org, and Blackmask.com and Bookshare.org.

I have not encountered any real stumbling blocks so far. There were a few things that took some time to figure out. For example, when my first etext was ready, I was pretty sure that it was expected that I would put the PG header on myself, but I looked all over the web site and could not find a "master" copy. (Actually, I think the master, such as it was/is, is available on Lyris, but I was not subscribing to Lyris then.) So I just pulled the header from a very-recently posted etext, but then after I sent the etext in it was posted with a different header anyway. (Nowadays, my understanding is that the PG "staff" prefers to put the header on.) I also spent some time researching 8-bit code pages, but I expect that the new big-FAQ will provide easy access to all the answers that I had to hunt down then. There's a lot of good information buried in past messages on the volunteers' board, but no good way to search out information on a particular topic.

So far I've been able to fill all my book needs without spending much money. I find my books through ABEbooks, or from Ebay, plus I've gotten a few at Ohio Book Store downtown on Main Street. I've rarely paid as much as $20 for a book, even including shipping. There's one book that I've purchased (but not yet started work on) which costs $1000 or more for the original edition, but which is also available in paperback reprints for about $10. There are some other books in my future plans which look like they will be more expensive, but we'll worry about that when the time comes.

My wife still cannot understand why I spend my time scanning books, whereas my kids (and, I guess, most other people I know) seem to think it's a little eccentric but basically acceptable behavior. Personally, I definitely enjoy producing etexts and hope to keep doing so for a long time. My thanks to Michael Hart, Jim Tinsley, Greg Newby, and untold others who devote so much effort to nurture the project and grease the skids for the rest of us. Long live Project Gutenberg.

Lynn Hill

I have been involved with PG since 1994, when I first began reading texts on-line during slow times at the office where I worked. (I once got into trouble with a co-worker when she found me "processing" Little Women instead of the week's payroll report.) I was surprised to find, even then, such a wide variety of material in the PG archives. I found myself re-reading favorite books from my childhood, and delighting in finding "new" ones--Little Lord Fauntleroy, The Secret Garden, Heidi, the Oz stories. They were not at all like the sugary old films I had seen on television. They were funny, heartwarming, and utterly charming. After some years as a reader of the texts, I found myself thinking, "I'd like to try this."

When I first checked out the web page for volunteers, I felt overwhelmed. There were all sorts of FAQ's, but when I read them, I was baffled by all the information about file types, fonts, and other details. I didn't even know where to get books, let alone what to do about jagged rights edges or indented lines. It was frustrating -- I had all this enthusiasm but didn't know where to apply it. I dawdled for some months, then came back and turned to the PG Volunteers' message board for help.

Help came from many sources. I found someone who needed a file proofread, so I offered to read it. This worked out well, and I even found a couple of typos in it. I proofed some more files for this person, and then some for other people on the board.

After a while, I was ready to try a whole book -- and from Dianne Bean came my first PG book, "The Golden Slipper" by Anna Katharine Green. When I opened the box, a stale smell floated out, and then I found a chunky book with the ugliest green cover I've ever seen on anything. The date was 1915, and the book was starting to crumble all around the edges. My first reaction was "Who would ever want to read this???" But since I had promised to do it, I dutifully started scanning and reading as I went along. The book was a collection of mystery/suspense stories about a teenage crime-stopper named Violet Strange. (I always felt as if Scooby Doo and his friends might turn up at any moment.) As I read, I began to like Violet, and to notice how different her world seemed from ours. By the time I reached the end of the book, I felt proud of myself for "saving" some good stories for the future, and ready to try another book.

My suggestion to new PG'ers is to jump in and not be shy about volunteering. PG is a big group of great people who care, but they do not know you are out there until you say something. Once you speak up, they will do anything short of triple backflips to help you.

There are many ways new folks can join in, from scavenging old books at yard sales all the way up to proofing files or scanning and typing in whole books. When you send in your first copy of title page and verso, be patient -- it takes time for your copyright research to be done. This is a great time to do proofing on-line at one of the distributed proofreading web sites.

I get my books from library sales, yard sales, friends I met on the PG Volunteer board, and even from elderly neighbors who wanted to lend me favorite books they have saved. When you want old books, tell everybody you know. They may come up with a lot of eligible books you wouldn't have expected.

Comments

Log in to leave a comment.

The Project Gutenberg FAQ 2002Chapter XI: Strait of Magellan.--Climate of the Southern Coasts (4)

0%37 min left in chapter