Chapter IV: Front Matter (4)
The rule of thumb is that we try not to post texts shorter than 25K, or about 350 lines of 70 characters. This rules out, for example, a lot of individual short poems. If you are interested in contributing this type of material, consider making a collection of similar texts--poems by the same author, or magazine articles on the same subject. We have made a few exceptions, like Martin Luther King's "I have a dream" speech, but very few.
V.18. What books are eligible?
A book is "eligible" for posting if we can legally publish it. This is the case if:
1. it is in the public domain in the U.S.A.,
OR,
2. the copyright holder has granted unlimited
non-exclusive distribution rights to PG.
V.19. Are reprints or facsimiles eligible?
A reprint or facsimile of a book that would be eligible is itself eligible.
For example, if a book published in 1995 is a reprint of a book published in 1900, then it is eligible. However, the onus is on us to prove that it _is_ a reprint, and if it doesn't _say_ on the TP&V that it is a reprint, confirming its eligibility may be impractical.
V.20. What is the difference between a reprint and a facsimile?
A facsimile retains the page layout and formatting of the original. A reprint keeps the same words, but may lay the pages out differently. For our copyright purposes, there is no difference--we can use either.
V.21. What is the difference between a reprint and a "new edition"?
A reprint contains only the words and pictures that were printed in the original. A new edition is in some way changed; it has different text, or pictures. It may be abridged, or expanded. It may have material added or changed, using other versions of the book.
A new edition gets a new copyright, and has to be cleared based on its own copyright date and status, not the date of the original printing of the title. See also the FAQ "How come my paper book of Shakespeare says it's 'Copyright 1988'?" [C.16] for an example.
Please note that we are talking here about a new edition of the printed book, not a new (corrected) edition number for Project Gutenberg naming purposes.
V.22. What book should I work on?
Nobody in Gutenberg is going to set assignments for you. You decide what book to process. Just pick one that no-one else has already done, or is working on. It's also sensible to pick one that you'll like--you'll be living with it for a while. On a practical note, it's probably better to start with a short book or even a short story, since a long book can take quite a while to produce.
Start by thinking of books written before 1923. Pick a book you like, and check it out. If it's already done or still in copyright, try other books by the same author.
Visit the Project Gutenberg site and download a full list of Gutenberg books in GUTINDEX.ALL. Have a look at the List of Books In Progress and Complete [B.1]. Look for authors you like, and see what books by them aren't yet available.
Check out your old books. Maybe you have an eligible edition that would be of great help to the project.
Try your library. They may have some eligible editions--books we can prove to be in the public domain--and you will certainly come away with ideas. Ask your librarian. Librarians are keen to help on projects like this.
Browse second-hand bookshops in your area. There are lots of treasures to be picked up very cheaply.
Search for literature pages and bookshops on the Internet.
If all else fails, you can always ask on the Volunteers' Board or try the gutvol-d mailing [V.12] list for ideas. Others may know of books that people are especially looking for, or projects already started where you could help out.
V.23. I have a book in mind, but I don't have an eligible copy.
First, determine whether there are any eligible copies of the book, by finding out the date it was published, possibly from the Catalog of the Library of Congress [B.5] and checking the Public Domain and Copyright Rules [B.1]. If there is a public domain edition, the next problem is to find one to work with.
V.24. Where can I find an eligible book?
The most commonly used outlets are used bookstores, garage sales, library sales, charity shops and any other place that sells old books.
The Internet is a wonderful medium for finding used and antiquarian books--used bookstores all over the world have found ways of co-operating and listing their inventories on the Net, so that whether you live in Los Angeles, Moscow or Perth, you can still find that book you're looking for in a shop in a laneway of Amsterdam. Most on-line listings will quote the publication year of the book, so you can check that it's pre-1923.
Two such sites that allow second-hand booksellers to list their inventory are:
Advanced Book Exchange <http://www.abebooks.com>
Alibris <http://www.alibris.com>
The book search page at trussel.com [B.5] has a list of many such Net bookshops, or you can simply visit any search engine and search for Used or Antiquarian Bookshops. You can often buy eligible books through these sites very cheaply.
If you still can't find the book you need, post a message on the Volunteers' Board or to the gutvol-d mailing list; maybe someone else can find it for you.
Sometimes, it may be possible for you to work from a later edition, so long as somebody who has an eligible edition can check it to make sure that no changes have been made. Sometimes, you may be able to find a modern reprint; reprints may be eligible, as long as they say they are reprints of an edition that would be eligible.
If you can type, or can scan without damaging the book, you can borrow books long enough to produce them. Even if your local library doesn't have the books you want, they may well be able to get them for you on inter-library loan. Ask your librarian about it.
V.25. What is "TP&V"?
This is an abbreviation for "Title Page and Verso", and means a paper or image copy of the front and back of the title page.
Even if the back is blank, we need to have an image of it for the files, to show that it _is_ blank, so that if, in ten years' time, somebody queries our right to publish, we can show that we haven't just lost it.
Publishers print copyright information, like title, author, copyright year and owner, and whether the book was a reprint, on the TP&V, and by filing this, we can prove that the book we produced was in the public domain.
Sending us the TP&V is the One True Way to getting PG copyright clearance [V.37].
V.26. What is "Posting"?
Posting is the final stage in the production process, where the file is given a number and official PG header, and copied onto our FTP servers for distribution. See section 4 of the FAQ "How does a text get produced?" [V.16] for a blow-by-blow account.
V.27. I think I've found an eligible book that I'd like to work on.
What do I do next?
Make sure nobody else is working on it, and that it's not already online somewhere.
V.28. What books are currently being worked on?
Check out David Price's In Progress List (a.k.a. "the InProg List") online at <http://www.dprice48.freeserve.co.uk/GutIP.html>. David gets the information from Copyright Clearances that have been done, and organizes it into a list. It can never be 100% up to date, since clearances come in all the time, but it's the best online facility we have, and it's much more clearly presented than the original clearance files.
V.29. How do I find out if my book is already on-line somewhere?
There's no foolproof method; some student somewhere could have scanned it and put it on her college web page without announcing it anywhere. However, there are some regular places to check.
It may sound obvious, but you should always look in the PG archives first. Download GUTINDEX.ALL and keep it handy. Search the InProg List [B.1].
The two other main places to search for your book are the Internet Public Library <http://www.ipl.org> and the On-Line Books Page <http://onlinebooks.library.upenn.edu/>. These projects specialize in indexing books that people make available on-line.
If you still don't see your book on-line anywhere, hit your favorite search engine, and give it the title, author's last name, and preferably a few uncommon words from the first page of the book. Sometimes one of those solo efforts shows up in a general search.
V.30. My book is not on the In-Progress list, and I can't find it on-line.
Is it safe to go ahead and buy it?
Probably. It could have been cleared, but not included in the InProg list yet. If the amount of money to buy it is a consideration, you can e-mail any of the members of the Posting Team, and ask them to check the latest clearances for you. Even this isn't foolproof; another volunteer could be placing their order at the same time you're placing yours. Such duplications do happen, but they are very rare.
V.31. My book is on-line, but not in Project Gutenberg. What should I do?
If the on-line file is from the same edition as the one you have (e.g. not a different translation) then you may be able to submit that file, perhaps slightly edited, to Gutenberg using the clearance from your paper copy. See "I've found an eligible text elsewhere on the Net, but it's not in the PG archives. Can I just submit it to PG?" [V.62] for how to do that.
And of course, you can always still make your own version for PG. It's surprising how often even very similar paper editions have small differences that can be interesting or significant.
V.32. My book is already on-line in Project Gutenberg, but my printed book
is different from the version already archived. Can I add my version?
Yes! In fact, assuming that the version already there is in the public domain, you can piggyback on the work already done by what is called "comparative retyping". For example, let's say that you have a later edition than the existing file; you can just take the existing file, edit it to match your paper version, and submit it as a new file. Of course, you must have Copyright Cleared [V.37] your paper version as well.
V.33. I see a book that was being worked on three years ago. Is anyone
still working on it?
Maybe, maybe not. Some people abandon books, some people who are regular producers clear them and put them at the bottom of the pile, perhaps for years (though they will get round to them sometime), and some people just simply take two or three years to produce a book.
Once, we put names and contact details on the public InProg list, but for privacy and spam-prevention reasons, we've taken them off. However, the Posting Team have access to the master list of cleared files, and will send a message on your behalf to the person who originally cleared the book, asking if the project is still active, or if the producer wants help.
So if you really want to check this situation out, e-mail one of the Posting Team.
V.34. I've decided which book to produce. How do I tell PG
I'm working on it?
As soon as you get Copyright Clearance [V.37], your book is entered in the "cleared" files. David Price will take these, and add your entry in his next release of the In Progress List.
V.35. I have a two- or three-volume set. Should I submit them as one
text, or one text for each volume?
Both.
Quite a lot of 18th and 19th Century books, even straightforward novels, were published as multipart sets. When you have such a set, you should usually submit one text for each volume, and a "complete" text with the contents of all volumes together.
People who do this often complete and submit one volume at a time, until they've finished, and then contribute the "complete" file.
V.36. I have one physical book, with multiple works in it (like a
collection of plays). Should I submit each text separately?
If the works are clearly separate, stand-alone texts, and are long enough [V.17] to warrant inclusion on their own in the archives, then yes, you should, and you _may_ also submit a "complete" version as well, if it seems appropriate. This most commonly happens in a collection of plays, though essays and other works may also fit the criteria. Collections of poetry rarely do, since most poems are too short to submit as stand-alone texts.
Sometimes the book includes a preface or introduction or glossary covering all the works in it. In this case, you can decide whether to include these with each of the parts, or save them for the "complete" version.
V.37. How do I get copyright clearance?
Basically we need to see images of the front and back of the title page of the book, which is where copyright information is usually shown. This is called "TP&V", for "Title Page and Verso" [V.25].
To Submit Online:
As of late 2002, we have a new automated upload procedure using a web page. This is by far the fastest and easiest way to get clearance. You need scanned images (PNG, JPEG, TIFF, GIF), of the two pages, of good enough resolution that the text can be read clearly, though the files don't need to be huge.
Just go to <http://beryl.ils.unc.edu/copy.html> and follow the instructions.
There are two other, older ways to submit a text for clearance.
To submit by paper mail, photocopy the front and back of the title page, even if the back is blank, write your e-mail address on it, and send the photocopies to:
MICHAEL STERN HART
405 WEST ELM STREET
URBANA, IL 61801-3231 USA
This is called Title Page & Verso, or TP&V for short, and is needed for copyright research. A colored envelope is best, to make sure your letter is easily recognized as TP&V.
E-mail Michael [email protected] when you send them, so he knows they're on the way. It's a good idea to check back with him by e-mail after a week or so if you haven't heard from him.
About this, Michael says: "Please include always your e-mail name and address, and mark the envelope with some distinctive mark and or color. Colored envelopes fine. Just something so I can find it easily, the mail here is slow and deep, like snow. Please send a note to: <[email protected]> for more info."
To submit by e-mail, scan the front and back of the title page, even if the back is blank, and e-mail the images to Greg Newby <[email protected]> as TIFF, JPEG or GIF in medium resolution. Make sure that the print is legible before you send.
Whichever method you use, you should expect to get an e-mail back after about a week, with one line containing the Author, Title, your name and date with the word "OK" at the end. This means that your text has been cleared.
A Clearance Line looks something like:
The Works Of Homer [Iliad/Odyssey] Tr. George Chapman Jim Tinsley 06/14/01 ok
If you don't get any response, e-mail to check that your TP&V was received OK. If the word at the end of the line is not "OK", then your text is not eligible, and a comment will probably be appended explaining why it is not eligible.
Don't start work on your book until you get that OK! It's very sickening to do all that work, and then find out that your text can't legally be put on-line!
V.38. I have a two- or three-volume set. Do I have to get a separate
clearance on each physical book?
Yes.
Some multi-volume works, notably reference books and translations, were published in a series, and it may be that the first volume is 1922, but the others are 1923 or later, so we have to clear each individually.
V.39. I have one physical book, with multiple works in it (like a
collection of plays). Do I have to get a separate clearance
for each work?
No. Since they were all printed together, one TP&V will suffice for all, but . . .
You should list each separate title included, if you intend to submit each title separately (see the FAQ "I have one physical book, with multiple works in it like a collection of plays. Should I submit each work separately?" [V.36]). If, say, you clear a "Collected Plays of Sheridan", and later submit an eBook of "The School for Scandal", we will have trouble finding your clearance unless we have made a note that "School for Scandal" is part of the contents of "Collected Plays".
In a case like this, you should include, on your paper or e-mail, something like:
George Bernard Shaw. Plays Unpleasant. 1905.
Contents:
Preface to Unpleasant Plays
Widower's Houses
The Philanderer
Mrs. Warren's Profession
You only need to do this when you are going to submit each part separately, which is commonly the case with plays, and sometimes essays, stories and novellas. Taking a different example, the "Collected Poems of Emily Dickinson", we would not need to list the contents, since we wouldn't publish each poem separately.
There is one exceptional case: if your book was printed after 1923, but contains stories or plays some of which are stated to be reprints of pre-1923 editions, you should give as much detail as possible about what you intend to submit.
V.40. Who will check up on my progress? When?
Nobody. There are no schedules or timetables. You're welcome to contact other volunteers [V.12] with comments or questions, though.
V.41. How long should it take me to complete a book?
Most books get done in between one and three months, but this varies wildly. It depends on the amount of time you can afford to give it, the length of the book and, if you're not typing, the quality of the scan--if the book scans badly, you need to put more time into proofing.
Some very productive volunteers manage to turn out an e-text a week; some books can take a year or more.
Scanning itself doesn't take too long. Even if it takes you as much as two minutes per page to scan, you will still complete a 300 page book in 10 hours, and you will probably be scanning much faster than that [S.9]. The problem is that the text generated by the scanner and your OCR package is usually faulty. There are many cute scanner errors, mistaking b for h, or e for c, so that "heard" is scanned as "beard" or "ear" as "car". Makes the story more interesting sometimes!
So now you need to do a first proof of the e-text. Read it carefully, correct scanning mistakes, and make sure that you haven't left out pages or got them in the wrong order. Unless your scan was exceptionally good, this is the time-burner in the process.
When you've done the first proof, you can either do a second proof yourself, or send it to another volunteer for second proofing.
If you're a typist, of course, you can skip right over the messy scanning and scan-correction process. Yay typists!!
V.42. I want/don't want my name published on my e-text
No problem. When you send the e-text for posting, mention exactly what, if anything, you want the Credits Line [V.47] to say.
V.43. I'd like to put a copy of my finished e-text, or another
Gutenberg text, on my own web page.
Great! PG encourages the widest possible distribution of e-texts. We like to publish everything in plain text, which is the most accessible format, since everybody can read plain text. But once it's available in plain text, it's open to you or anyone else to convert it to other formats like HTML for further distribution.
If you are reposting a text, though, please be careful to check that your posting complies with the conditions spelled out in the header, especially for copyrighted works.
V.44. I've scanned, edited and proofed my text. How do I find someone
to second-proof it?
You can post a request on the Volunteers' Board, or on the gutvol-d Mailing List. You will probably get some offers there. In a difficult case, you might ask Michael Hart to add it to the "Requests for Assistance" section of the next Newsletter.
In general, the best way to handle it is to make a co-operative proofing project out of it. This is like a miniature version of the distributed proofreading sites, without the page images.
There are always people looking for proofing work, but many beginners take on more than they can handle, and don't finish the job, and this can be very disappointing if you give the whole thing to one volunteer who then vanishes without trace. You can minimize the risk of this by splitting the book into chunks of about 20-30 pages, or one chapter if that's around the right size, each. Write explicit instructions about what you want them to do when they spot a suspected error, like fix it or mark it with an asterisk. (Marking is probably safer with beginners who don't have the book or an image of the page to refer to.) Give the first chapter to the first person who responds, the second to the second, and so on. As you hand out the chapters, let the proofers know that if they're not returned within three or five days, you'll assume they've quit. Three days is more than plenty of time for 20 pages. If someone returns a chapter, you can give them another. If someone doesn't get back to you within the time set, assume they're not going to, and recycle that chapter to someone else. No hard feelings, no problem. This process of "co-operative proofing" ensures that beginning proofers don't choke on the work, and that one vanishing volunteer doesn't hold up the whole project.
V.45. I've gone over and over my text. I can't find any more errors,
and I'm sick of looking at it. What should I do now?
We all know that feeling! Particularly with your first book, you've probably gone through a patch when you thought you'd never finish--and when you do, you can't stand the idea of looking at it again. Heh. Cheer up--the first twenty texts are the worst! :-) And you'll feel a lot better when you see your text available for everyone to read.
You have three choices:
You can send it for posting as it is. [V.46]
You can put it aside for week or so, and come back to it with fresh eyes.
You can ask in any of the standard ways [V.12] for someone else to second-proof it for you. This has a lot to recommend it; it gets other sets of eyes looking at the text, it relieves the pressure that you may feel, it may rekindle your enthusiasm for the text, it allows you to "meet" other volunteers, and possibly form partnerships for future PG collaboration. Above all, it gives new proofers a chance to get their feet wet, and this is good for them, and good for PG. You are not only contributing a text, you're helping to train and encourage the next generation of producers.
V.46. Where and how can I send my text for posting?
As of late 2002, we have a new automated upload procedure using a web page. This has a lot of good things going for it, because we keep a record of what's uploaded, you get an e-mailed copy of the notification, you don't have to fiddle with FTP, and we can make up the header automatically from the information you enter, which saves time and prevents keying errors.
As always, it's better to ZIP your file first, because it'll take less time to transfer.
Just go to <http://beryl.ils.unc.edu/cgi-bin/upload>, fill in the form, specify the file to upload, and hit "Send" at the bottom.
And you're done!
If, for some reason, you can't use this page, there are two backup options: you can e-mail it, or you can upload it by FTP. Whichever you use, it is always best to ZIP the file first if you can.
If you are comfortable with sending files by FTP, this is better than e-mail, First, you will need a username and password, which you can get by e-mailing any of the Posting Team.
If you already know how to use command-line FTP, here's how to do it:
Log in to beryl.ils.unc.edu using the username and password supplied and change to the work directory by typing "cd work". Change to binary mode with the "bin" command and "put" your file.
Summary instructions: ftp beryl.ils.unc.edu login: yourlogin password: yourpassword cd work bin put yourfile.ext quit
Here is a sample session:
>ftp beryl.ils.unc.edu Connected to beryl.ils.unc.edu. 220-Access from [email protected] logged. 220 FTP Server User (beryl.ils.unc.edu:(none)): xxxxxxxx 331 Password required for xxxxxxxx. Password: xxxxxxxx 230 User xxxxxxxx logged in. ftp> cd work 250 CWD command successful. ftp> bin 200 Type set to I. ftp> put MYFILE.ZIP 200 PORT command successful. 150 Opening BINARY mode data connection for MYFILE.ZIP. 226 Transfer complete. ftp: 172313 bytes sent in 17.34Seconds 9.94Kbytes/sec. ftp> quit
When you are in the work directory, you will not be able to list files, but they _do_ exist and they _are_ there.
When you have uploaded your file, e-mail a note to any or all of the Posting Team, including your 1. filename 2. credits line as you want it on your text 3. clearance line you received [V.37]
An ideal note might be:
Subject: Beryl upload for posting: Hamlet
I have uploaded to beryl:
Hamlet, by William Shakespeare
File is: hamlet.zip
Credits line is:
Produced by John Doe <[email protected]>
Clearance was given as:
Hamlet William Shakespeare John Doe 05/03/02 ok
If you'd rather send it by e-mail, send the e-mail, including the Credits Line and Clearance Line as in the sample above, to any or all of the Posting Team, with your text as an attachment. Again, ZIPped is better, since it avoids certain damage that can happen to a plain text e-mail along the way.
Do not add the Project Gutenberg header or footer to your file, unless we specifically asked you to. If you do add it, we'll just have to strip it off again, since we add headers automatically when posting. There are times, perhaps when you're working in an unusual non-editable format, when we may give you a header and ask you to add it, but this is rare.
Please read section "4: Posting" of the FAQ "How does a text get produced?" [V.16] for more detail about what happens in posting. Especially, if you want to draw some peculiarities of this text to the Posting Team's attention, or want feedback on any minor edits done during posting, you should say so in the e-mail you send.
_Don't assume that we know anything_ when you send the e-mail. We don't know what you want us to put on the Credits Line. We don't know that this is an unusual text, and needs some kind of special reformatting. We don't know that the text should be split into two volumes before posting. We don't know that you would really like us to check it closely before posting. You have to tell us, exactly and precisely, what you want on the Credits Line. If the text needs some specific work, you have to tell us exactly what that is. And please do that in your e-mail, not in the text itself. Remember that we could be dealing with five or ten other texts at the same time, and even if the poster you discussed it with two weeks ago is the same one who posts the book, he may not remember.
V.47. What is the "Credits Line"?
The Credits line is a line that the Posting Team can insert into each PG text naming the producer or producers of a particular text.
You should decide what you want on the credits line of your text; it's really not up to us.
Most credits lines are something like:
Produced by John Doe <[email protected]>.
If you don't want to be mentioned by name at all, just say, in your e-mail:
Please omit the Credits Line for this text. I want to contribute
it anonymously.
If you do want to be mentioned, please give the exact wording you want us to use. Some people want their name only; they don't want us to include their e-mail addresses. Others want to make their e-mail addresses public so that readers can contact them with comments. That is entirely up to you, but you do need to tell us. If you do want to include your e-mail, remember that having it permanently on the net is a spam-magnet, and we can't effectively remove or change it later.
Occasionally, a Credits Line can spill onto more than one line, for example:
This text was converted to HTML by Jane Roe <[email protected]>
from an original ASCII text scanned by Jack Went
and proofed by Jill Hill
V.48. How soon after I send it will my text be posted?
First read the "Posting" section of the FAQ "How does a book get produced?" [V.16] to understand the process.
You should expect some response within three or four days. We try to get to all submissions within that time. In most cases, that response will be simply the official notification that it has been posted. If there is a query on your text, for example if we can't find the copyright clearance or if we have trouble converting or correcting your text, we will probably e-mail you back directly with questions.
If you don't hear from us within four days, send a follow-up e-mail; it could be that your original note never got to us, or just fell through the cracks.
If your file happens to arrive while one of us is logged in and working, it could get posted within the hour. Some frequent contributors who know our habits know just how to time their uploads!
V.49. I found a problem with my posted text. What do I do?
Most postings go smoothly, but problems can happen. Sometimes, one of the servers is down. Sometimes a file gets corrupted for some unknown reason. Sometimes, let's face it, we screw up.
Usually, one of the indexers will tell us about it, but if you catch it first, e-mail whoever sent out your notification e-mail and explain the problem. Don't worry; your original file will be quite safe, since we keep these long after posting them.
V.50. Someone has e-mailed me about my posted text, pointing out errors.
Great!
Since you're the original producer, you're in the best position to decide whether these are real errors. If they're right about it, tell the Posting Team and we'll correct the text.
V.51. Someone has e-mailed me about my posted text, thanking me.
Nice feeling, isn't it? :-)
About Proofing
V.52. What role does proofing play in Project Gutenberg?
A very big one!
Typists' work doesn't usually need many corrections, but unfortunately, scanners and OCR packages are far from perfect, and scanned text varies from "almost-right" down to "maybe I should consider typing instead of scanning". Proofing is the process that turns a scan into a readable e-text.
Proofing a typist's work is straightforward; you just read it, and keep an eye out for mistakes. Typists typically have few mistakes in their texts, but the errors that they do make tend to be hard to spot. Proofing OCRed text has its quirks, and you can expect many, many errors to correct.
The only thing that all proofers agree on is to differ in their methods. Some people scan and almost complete the proofing process within their OCR package, others do no editing at all within their OCR. Some spell-check first, others spell-check last. Some work through in one pass, doggedly line by line, others make several light passes. Some start at the end and work backwards! Some proofers mark all queries with special characters like asterisks (*) in the text, most just make all the obvious changes and mark only the dubious ones. Some people always send their texts out for proofing, others prefer to do it all themselves.
So this guide is not prescriptive; this is not the "only way" to do it. The only rule is that, at the end of the process, your e-text should be as error-free as you can make it, and should conform to Gutenberg's editing standards, which are mostly just common sense guidelines to make readable text.
The aim of this FAQ is to give you an understanding of what text looks like when it comes fresh off the scanner, and an overview of the whole process by which it becomes a publishable e-text.
V.53. What is Distributed Proofing?
It has always been common for volunteers to share proofing work among themselves--you take the first five chapters, I'll take the next, and so on.
When you're just starting as a PG volunteer, you should go to one of the Distributed Proofing sites [B.4] and do some work there to get a grounding in the basics and a feel for whether you would like to continue working in PG. In distributed proofing, you get a very short section, as little as a page of text at a time, and usually an image file of the page as it scanned. You then make the text match the image. This is a great start, since all you have to do is read, compare and correct. However, other work also needs to be done, and will normally be done by the project managers of these sites. The samples below give you an idea of the whole process, and also some ideas of what proofing a whole book from start to finish is like.
V.54. What do I need to proof an e-text?
You actually need only two things: the e-text itself and a text editor or word-processor that can handle book-sized files and save them as text.
Nearly all word processors and text editors in current use will work. Volunteers use many common programs, including WordPerfect, Microsoft Word, WordPad, DOS EDIT, vi, Brief, Crisp, EditPad, MetaPad, emacs, AbiWord, and the word processors from Open Office abd AppleWorks. And all of these are in actual use by volunteers today. Since all of them contain the necessary basic functions, the best program is the one you're most comfortable with.
Be cautious with recent, powerful word-processors that "auto-correct" text, or use "smart quotes" or any other such automatic retyping or formatting feature, since they can Do Bad Things to your e-text without your consent! When using any such package, it is best to switch off any feature that makes changes without asking you.
Two utilities which may come in useful are a spell-checker and a version difference checker. These may be built into your word processor, or you may have them as separate packages.
A spell-checker is like a chain-saw: a powerful tool, but one to be used very carefully. It is very easy to say "Yes" to the wrong change, and make a really bad mess of the text. Spell-checkers have problems with proper names, foreign words, archaic usages, and dialects. Incautious use can leave you with a text such as that immortalized in the
Owed two a Spell in Chequer.
Eye half a spell in chequer,
It cane with my Pea Sea.
It plane lee marques four my revue
Miss steaks eye can knot sea.
Every e-text should pass through a spell-checker at some point, but the human half of the partnership needs a very light hand on the confirmations of change!
A difference checker, such as FC or COMP for MS-DOS, diff for Unix or ExamDiff <http://www.prestosoft.com/examdiff/examdiff.htm> for Windows, may also come in handy. A difference checker compares two versions of the text, and points out the changes. This is important when you've sent a text out for proofing, and you get it back with changes. Rather than re-reading the whole text, you can use a difference checker to highlight the changes so that you can verify them against the printed text. As a proofer, you can use it to compare the original text with what you're sending back to ensure that you've only changed what you meant to change.
V.55. Do I need to have a paper copy of the book I'm proofing?
No.
Your job as proofer is to ensure that the e-text you're working on is readable in itself, and contains no obvious errors. Where you think there might be an error, but you're not sure, you mark the spot in the e-text, and let the volunteer who has the paper book look it up.
V.56. What's the difference between "first proof" and "second proof"?
These are fuzzy terms used to indicate how accurate the e-text is, and what type of work is needed to improve it. Quite commonly, the same volunteer who scans the book proofs the whole thing in one or two passes. Sometimes, given a good scan, the text can be sent out for "first proof" with little or no preparatory fixing-up. Often, the scanner makes quite a lot of corrections, then sends the text out for "second proof".
A text is ready for first proofing when it's obvious that there are plenty of errors, but it's possible to figure out, in almost every case, what the correct text should be without needing to refer to the book.
The objective of first proofing is to eliminate all the obvious errors, so that if you speed-read quickly through the text, you probably won't notice any.
Second proofing involves taking a text that has been first-proofed and correcting all the remaining, more subtle errors. Often, some simple errors such as incorrect spacing and quotes may be left for second proofing. Texts that have been typed instead of scanned will always be of at least second-proof quality.
V.57. What do I do with an e-text sent to me for proofing?
First, establish reasonable expectations. A typical book takes 10-15 hours of concentrated effort, and when you first start, you're climbing a learning curve. For your first session, decide to mark out a chapter or two--something like 500 to 1,000 lines--and work only on that. If you get through 1,000 lines in your first sitting, you have done extremely well! It's a good idea to send this first 1,000 lines or so back immediately. The volunteer who sent you the e-text will comment on it, and let you know about any style guidelines you may have breached or common errors you may have missed. Most beginning proofers do make mistakes, so don't worry about it--it's easier to correct these in 1,000 lines than to go back over them in 15,000 lines!
You will usually receive the e-text as an attachment to your e-mail. It's better to send e-texts as attachments than to paste them as text into the body of the e-mail to make sure that the text isn't changed by different e-mail clients. It's better to send e-mailed attachments as ZIP files [R.20], since e-mails sent as text can be damaged along the way. But whether you receive a TXT file or a ZIP file that you have to open, you should save the .TXT file to your hard disk and open it with your editor.
It may be that the text you see appears double-spaced--every second line is blank--or that all the text is on one incredibly long line. This is a familiar effect when moving between a DOS/Windows computer and a Mac or Unix system, but it can happen between any two editors. It is caused by the use of different characters to mark the end of a line. If you have this problem, ask whoever sent you the text to re-send it, telling them what kind of computer and editor you have.
Now you make any changes that obviously need to be made, and mark any places where the text looks wrong, but you're not sure what the right text should be. You can usually use asterisks (*) to mark these dubious spots, but you might use other characters if the text already contains asterisks. When in doubt, mark them all, and let the volunteer with the text sort them out!
It is usually best not to make global changes to line lengths by reformatting lots of paragraphs, since the person who sent you the e-text may want to use a difference checker when you return it, and changed line-lengths throughout mean that every line will be different.
When working on a long text, or when making a lot of changes, it may be wise to save several versions of the text with different filenames at different stages so that if something goes badly wrong, you can revert to the last good version. This applies especially to saving the text just before performing a spell-check.
When you're finished with the e-text, make sure you save it as a plain text file (.TXT) and send it back by zipping it if you can, and attaching it to an e-mail.
V.58. What kinds of errors will I have to correct?
Each text has its own peculiarities, but there are a number of well-known scanning errors you will be dealing with all the time.
Punctuation is always a problem. Periods, commas and semi-colons are often confused, as are colons and semi-colons. There are also usually a number of extra or missing spaces in the e-text.
The problem of quotes can assume nightmarish proportions in a text which contains a lot of dialog, particularly when single and double quotes are nested.
The numeral 1, the lower-case letter l, the exclamation mark ! and the capital I are routinely confused, and often, single or double quotes may be mistaken for one of these.
Lower-case m is often mistaken for rn or ni.
The letters h and b and e and c are commonly mis-read, and these are probably the hardest of all to catch, since ear/car, eat/cat, he/be, hear/bear, heard/beard are all common words which no spell-checker will flag as problems.
For example:
" Hello1' caIled jirnmy breczily. 11Anyone home ? "
There seemed to he no-oneabout. Only tbe eat beard him."
should read:
"Hello!" called Jimmy breezily, "Anyone home?"
There seemed to be no-one about. Only the cat heard him.
As well as scanner errors, which affect one letter at a time, you have to keep an eye out for editing mistakes by the volunteer who scanned the text or by previous proofers. These are typically cases where a whole line, paragraph or page has been omitted or misplaced. They show up as sentences that don't make sense, or paragraphs that don't follow from the previous one.
This means that you have to keep reading the flow of the text, so that you can spot context errors as well as typos.
V.59. How long does it take to proof an e-text?
This depends on how long the e-text is, how clean the text is when you start, and how thorough you're being, as well as how much time per day you can give it and how fast you can proof.
On a first proof, it can take a very long time to get the e-text to a readable condition if it scanned badly. As a beginner, you would be unlikely to be given such a difficult text to work with. First proofs are usually done by the same person who did the scanning, and are only given out in the context of established scanning/proofing teams.
You might expect to proof anywhere between 500 and 2,000 lines per hour during a second proof. A short novel or novella might have as few as 6,000 or 7,000 lines; War and Peace weighs in at about 54,000 lines. Most novels run to 10,000 to 15,000 lines. So you might spend anything between 5 and 30 hours second-proofing a standard book, with 10 to 15 hours being typical.
For an average novel, a week or two for second proofing is good going. A month is reasonable.
Proofing an e-text is a significant amount of work, and you may find it psychologically more comfortable to take on a chunk at a time--say 1,000 lines per session--and send that proofed section back, rather than wait until the whole job is done before sending anything back. This helps to avoid the fairly common case where you keep falling behind where you expect to be until you dread the thought of getting back to the text, and finally just abandon it.
If you find after a while that you just don't want to continue, please tell the person who sent you the text that you're not going ahead with it. It's very frustrating for the volunteer who scanned the book, and who wants to get it posted, to wait for two or three months, only to have to start all over again with another proofer.
V.60. Are there any special techniques for proofing?
The classic way to proof is to open the text in your editor or word processor, and just start reading carefully.
This method has received a major boost since editors and word processors have added a feature of showing squiggly red underlines under words not in their dictionary. While this is very useful, you still need to read carefully, since not all errors produce misspelled words. The classic, and very common, example of this is scanning "he" for "be". These visual spellchecks also commonly do not check words beginning with capitals. Capitalized words are commonly names not in the dictionary, and when checking of capitalized words is switched off, they will not query "Tbe". Other errors that a spellchecker doesn't look for include missing spaces, mismatched quotes and misplaced punctuation. For these, you can try gutcheck [P.1]. And of course, no automatic check will find omitted lines or words. Worse, spellcheckers will query words not in their dictionary that might be quite correct, and this can be quite troublesome when dealing with older texts or dialect.
Still, if your concentration is up to the job, scrolling through a text with non-dictionary words underlined in red is a fast and effective way of giving a text the final once-over.
Volunteers have also used other techniques for proofing. Some people can't sit at their screen and read for hours; many people don't want to.
Some people just use the good old-fashioned method of printing out the text to be proofed, and blue-pencilling the mistakes.
It is becoming fairly common now for people to load the text onto their PDA, and read it from that. Mistakes found can be bookmarked or jotted down and fixed when they go back to their PC.
Getting your computer to read the text aloud is a very effective way of achieving high accuracy. Modern PCs have audio capabilities built in, and it is possible to find free or cheap shareware "read-aloud" text-to-speech packages for just about everything. Some PDAs are also capable of doing text-to-speech.
The first time you try text-to-speech, it will probably sound and feel a little strange, but you will quickly learn to _hear_ errors in words. This can be very effective, but you should have given the text at least a light proofing before you begin; it is hard to deal with a high number of errors using a text-to-speech method.
When proofing by a speech program, you either set your text-to-speech program to pronounce all punctuation, or, if that is not possible, you make a special version of your text to feed it, first doing a global replace of "," with " comma ", ";" with " semi-colon ", and so on. Mark a block of 500 to 1,000 lines for reading aloud, and set the reading speed to whatever is comfortable for you. Then you sit down with the original book in front of you, and listen. When you hear an error, mark the place in the text with a light pencil. Stopping the reading at every error, editing the text and restarting is possible, but it breaks the flow, and ends up taking longer. When the reading is done, go to your keyboard and correct the errors found.
V.61. What actually happens during a proof?
Stage One--The original Scan
We start with a scanned e-text, in this case a paragraph from The Odyssey. The paragraph used as an example here has been "enhanced" with more errors than in the real scanned text, so that you can see samples of many problems all in one place.
We begin by looking at the original OCRed text, of which our sample section reads:
1There Periniedes and Eurylochus held the victims, but l
drew my sharp sword from my thigh, and dug a pit, as it were
a cubit in length and breadth, and about it poured a drink-
offering to all the dead, first with mead and thereafter with
sweet wine, and for the third time with water, And 1 sprink-
Comments
Log in to leave a comment.
The Project Gutenberg FAQ 2002Chapter IV: Front Matter (4)
0%34 min left in chapter