Improvement, kindle and README (thanks @nrenzoni)
This commit is contained in:
parent
2a7df72bd1
commit
b9d78bcd63
3 changed files with 60 additions and 53 deletions
3
.gitignore
vendored
3
.gitignore
vendored
|
|
@ -1,5 +1,7 @@
|
|||
.idea/
|
||||
.vscode/
|
||||
|
||||
__pycache__/
|
||||
.venv/
|
||||
.env/
|
||||
|
||||
|
|
@ -7,3 +9,4 @@ Books/
|
|||
|
||||
cookies.json
|
||||
*.log
|
||||
*.txt
|
||||
|
|
|
|||
78
README.md
78
README.md
|
|
@ -7,9 +7,10 @@ Before any usage please read the *O'Reilly*'s [Terms of Service](https://learnin
|
|||
* [Requirements & Setup](#requirements--setup)
|
||||
* [Usage](#usage)
|
||||
* [Example: Download *Test-Driven Development with Python, 2nd Edition*](#download-test-driven-development-with-python-2nd-edition)
|
||||
* [Example: Use or not the `--no-kindle` option](#use-or-not-the---no-kindle-option)
|
||||
* [Example: Use or not the `--kindle` option](#use-or-not-the---kindle-option)
|
||||
|
||||
## Requirements & Setup:
|
||||
First of all, it requires `python3` and `pip3` or `pipenv` to be installed.
|
||||
```shell
|
||||
$ git clone https://github.com/lorenzodifuccia/safaribooks.git
|
||||
Cloning into 'safaribooks'...
|
||||
|
|
@ -22,7 +23,7 @@ OR
|
|||
$ pipenv install && pipenv shell
|
||||
```
|
||||
|
||||
The program depends of only two **Python 3** modules:
|
||||
The program depends of only two **Python _3_** modules:
|
||||
```python3
|
||||
lxml>=4.1.1
|
||||
requests>=2.20.0
|
||||
|
|
@ -44,30 +45,30 @@ Like: `https://www.safaribooksonline.com/library/view/test-driven-development-wi
|
|||
#### Program options:
|
||||
```shell
|
||||
$ python3 safaribooks.py --help
|
||||
usage: safaribooks.py [--cred <EMAIL:PASS> | --login] [--no-cookies] [--no-kindle]
|
||||
[--preserve-log] [--help]
|
||||
usage: safaribooks.py [--cred <EMAIL:PASS> | --login] [--no-cookies]
|
||||
[--kindle] [--preserve-log] [--help]
|
||||
<BOOK ID>
|
||||
|
||||
Download and generate EPUB of your favorite books from Safari Books Online.
|
||||
Download and generate an EPUB of your favorite books from Safari Books Online.
|
||||
|
||||
positional arguments:
|
||||
<BOOK ID> Book digits ID that you want to download.
|
||||
You can find it in the URL (X-es):
|
||||
`https://www.safaribooksonline.com/library/view/book-
|
||||
<BOOK ID> Book digits ID that you want to download. You can find
|
||||
it in the URL (X-es):
|
||||
`https://learning.oreilly.com/library/view/book-
|
||||
name/XXXXXXXXXXXXX/`
|
||||
|
||||
optional arguments:
|
||||
--cred <EMAIL:PASS> Credentials used to perform the auth login on Safari
|
||||
Books Online.
|
||||
Es. ` --cred "account_mail@mail.com:password01" `.
|
||||
Books Online. Es. ` --cred
|
||||
"account_mail@mail.com:password01" `.
|
||||
--login Prompt for credentials used to perform the auth login
|
||||
on Safari Books Online.
|
||||
--no-cookies Prevent your session data to be saved into
|
||||
`cookies.json` file.
|
||||
--no-kindle Remove some CSS rules that block overflow on `table`
|
||||
and `pre` elements. Use this option if you're not going
|
||||
to export the EPUB to E-Readers like Amazon Kindle.
|
||||
--preserve-log Leave the `info_XXXXXXXXXXXXX.log` file even if there
|
||||
--kindle Add some CSS rules that block overflow on `table` and
|
||||
`pre` elements. Use this option if you're going to
|
||||
export the EPUB to E-Readers like Amazon Kindle.
|
||||
--preserve-log Leave the `info_XXXXXXXXXXXXX.log` file even if there
|
||||
isn't any error.
|
||||
--help Show this help message.
|
||||
```
|
||||
|
|
@ -77,21 +78,26 @@ The next times you'll download a book, before session expires, you can omit the
|
|||
For **SSO**, please use the `sso_cookies.py` program in order to create the `cookies.json` file from the SSO cookies retrieved by your browser session (please follow [`these steps`](/../../issues/150#issuecomment-555423085)).
|
||||
|
||||
Pay attention if you use a shared PC, because everyone that has access to your files can steal your session.
|
||||
If you don't want to cache the cookies, just use the `--no-cookies` option and provide all time your `--cred` to perform `--login`.
|
||||
If you don't want to cache the cookies, just use the `--no-cookies` option and provide all time your credential through the `--cred` option or the more safe `--login` one: this will prompt you for credential during the script execution.
|
||||
|
||||
You can configure proxies by setting on your system the environment variable `HTTPS_PROXY`.
|
||||
You can configure proxies by setting on your system the environment variable `HTTPS_PROXY` or using the `USE_PROXY` directive into the script.
|
||||
|
||||
The program default options are thought for ensure best compatibilities for who want to export the `EPUB` to E-Readers like Amazon Kindle. If you want to do it, I suggest you to convert the `EPUB` to `AZW3` with [Calibre](https://calibre-ebook.com/).
|
||||
You can also convert the book to `MOBI` and if you'll do it with Calibre be sure to select `Ignore margins` in the conversion options:
|
||||
**Important**: since the script only download HTML pages and create a row EPUB, many of the CSS and XML/HTML directives are wrong for an E-Reader. To ensure best quality of the output, I suggest you to always convert the `EPUB` obtained by the script to standard-`EPUB` with [Calibre](https://calibre-ebook.com/).
|
||||
You can also use the command-line version of Calibre with `ebook-convert`, e.g.:
|
||||
```bash
|
||||
$ ebook-convert "XXXX/safaribooks/Books/Test-Driven Development with Python 2nd Edition (9781491958698)/9781491958698.epub" "XXXX/safaribooks/Books/Test-Driven Development with Python 2nd Edition (9781491958698)/9781491958698_CLEAR.epub"
|
||||
```
|
||||
After the execution, you can read the `9781491958698_CLEAR.epub` in every E-Reader and delete all other files.
|
||||
|
||||
The program offers also an option to ensure best compatibilities for who wants to export the `EPUB` to E-Readers like Amazon Kindle: `--kindle`, it blocks overflow on `table` and `pre` elements (see [example](#use-or-not-the---kindle-option)).
|
||||
In this case, I suggest you to convert the `EPUB` to `AZW3` with Calibre or to `MOBI`, remember in this case to select `Ignore margins` in the conversion options:
|
||||
|
||||

|
||||
|
||||
In the other hand, if you're not going to export the `EPUB`, you can use the `--no-kindle` option to remove the CSS that blocks overflow on `table` and `pre` elements, see below in the examples.
|
||||
|
||||
## Examples:
|
||||
* ## Download [Test-Driven Development with Python, 2nd Edition](https://www.safaribooksonline.com/library/view/test-driven-development-with/9781491958698/):
|
||||
```shell
|
||||
$ python3 safaribooks.py --cred "XXXX@gmail.com:XXXXX" 9781491958698
|
||||
$ python3 safaribooks.py --cred "my_email@gmail.com:MyPassword1!" 9781491958698
|
||||
|
||||
____ ___ _
|
||||
/ __/__ _/ _/__ _____(_)
|
||||
|
|
@ -118,32 +124,36 @@ In the other hand, if you're not going to export the `EPUB`, you can use the `--
|
|||
that works.In the process, you’ll learn the basics of Django, Selenium, Git,
|
||||
jQuery, and Mock, along with curre...
|
||||
[-] Release Date: 2017-08-18
|
||||
[-] URL: https://www.safaribooksonline.com/library/view/test-driven-development-with/9781491958698/
|
||||
[*] Retrieving book chapters...
|
||||
[-] URL: https://learning.oreilly.com/library/view/test-driven-development-with/9781491958698/
|
||||
[*] Retrieving book chapters...
|
||||
[*] Output directory:
|
||||
/XXXX/XXXX/Books/Test-Driven Development with Python, 2nd Edition
|
||||
[-] Downloading book contents... (73 chapters)
|
||||
[#########################################----------------------------] 60%
|
||||
...
|
||||
/XXXX/safaribooks/Books/Test-Driven Development with Python 2nd Edition (9781491958698)
|
||||
[-] Downloading book contents... (53 chapters)
|
||||
[#####################################################################] 100%
|
||||
[-] Downloading book CSSs... (2 files)
|
||||
[#####################################################################] 100%
|
||||
[-] Downloading book images... (142 files)
|
||||
[#####################################################################] 100%
|
||||
[-] Creating EPUB file...
|
||||
[*] Done: Test-Driven Development with Python, 2nd Edition.epub
|
||||
|
||||
[*] Done: /XXXX/safaribooks/Books/Test-Driven Development with Python 2nd Edition
|
||||
(9781491958698)/9781491958698.epub
|
||||
|
||||
If you like it, please * this project on GitHub to make it known:
|
||||
https://github.com/lorenzodifuccia/safaribooks
|
||||
e don't forget to renew your Safari Books Online subscription:
|
||||
https://www.safaribooksonline.com/signup/
|
||||
|
||||
https://learning.oreilly.com
|
||||
|
||||
[!] Bye!!
|
||||
```
|
||||
The result will be (opening the `EPUB` file with Calibre):
|
||||
|
||||

|
||||
|
||||
* ## Use or not the `--no-kindle` option:
|
||||
* ## Use or not the `--kindle` option:
|
||||
```bash
|
||||
$ python3 safaribooks.py --no-kindle 9781491958698
|
||||
$ python3 safaribooks.py --kindle 9781491958698
|
||||
```
|
||||
On the left book created with `--no-kindle` option, on the right without (default):
|
||||
On the right, the book created with `--kindle` option, on the left without (default):
|
||||
|
||||

|
||||
|
||||
|
|
|
|||
|
|
@ -239,11 +239,10 @@ class SafariBooks:
|
|||
"<head>\n" \
|
||||
"{0}\n" \
|
||||
"<style type=\"text/css\">" \
|
||||
"body{{margin:1em;}}" \
|
||||
"body{{margin:1em;background-color:transparent!important;}}" \
|
||||
"#sbo-rt-content *{{text-indent:0pt!important;}}#sbo-rt-content .bq{{margin-right:1em!important;}}"
|
||||
|
||||
KINDLE_HTML = "body{{background-color:transparent!important;}}" \
|
||||
"#sbo-rt-content *{{word-wrap:break-word!important;" \
|
||||
KINDLE_HTML = "#sbo-rt-content *{{word-wrap:break-word!important;" \
|
||||
"word-break:break-word!important;}}#sbo-rt-content table,#sbo-rt-content pre" \
|
||||
"{{overflow-x:unset!important;overflow:unset!important;" \
|
||||
"overflow-y:unset!important;white-space:pre-wrap!important;}}"
|
||||
|
|
@ -349,8 +348,6 @@ class SafariBooks:
|
|||
self.display.info("Retrieving book chapters...")
|
||||
self.book_chapters = self.get_book_chapters()
|
||||
|
||||
self.images = self.extract_image_links(self.book_chapters)
|
||||
|
||||
self.chapters_queue = self.book_chapters[:]
|
||||
|
||||
if len(self.book_chapters) > sys.getrecursionlimit():
|
||||
|
|
@ -376,9 +373,10 @@ class SafariBooks:
|
|||
self.filename = ""
|
||||
self.chapter_stylesheets = []
|
||||
self.css = []
|
||||
self.images = []
|
||||
|
||||
self.display.info("Downloading book contents... (%s chapters)" % len(self.book_chapters), state=True)
|
||||
self.BASE_HTML = self.BASE_01_HTML + (self.KINDLE_HTML if not args.no_kindle else "") + self.BASE_02_HTML
|
||||
self.BASE_HTML = self.BASE_01_HTML + (self.KINDLE_HTML if not args.kindle else "") + self.BASE_02_HTML
|
||||
|
||||
self.cover = False
|
||||
self.get()
|
||||
|
|
@ -613,7 +611,7 @@ class SafariBooks:
|
|||
|
||||
@staticmethod
|
||||
def is_image_link(url: str):
|
||||
return pathlib.Path(url).suffix[1:] in ["jpg", "peg", "png", "gif"]
|
||||
return pathlib.Path(url).suffix[1:].lower() in ["jpg", "jpeg", "png", "gif"]
|
||||
|
||||
def link_replace(self, link):
|
||||
if link and not link.startswith("mailto"):
|
||||
|
|
@ -814,6 +812,11 @@ class SafariBooks:
|
|||
self.chapter_title = next_chapter["title"]
|
||||
self.filename = next_chapter["filename"]
|
||||
|
||||
# Images
|
||||
if "images" in next_chapter and len(next_chapter["images"]):
|
||||
self.images.extend(urljoin(next_chapter['asset_base_url'], img_url)
|
||||
for img_url in next_chapter['images'])
|
||||
|
||||
# Stylesheets
|
||||
self.chapter_stylesheets = []
|
||||
if "stylesheets" in next_chapter and len(next_chapter["stylesheets"]):
|
||||
|
|
@ -1045,14 +1048,6 @@ class SafariBooks:
|
|||
shutil.make_archive(zip_file, 'zip', self.BOOK_PATH)
|
||||
os.rename(zip_file + ".zip", os.path.join(self.BOOK_PATH, self.book_id) + ".epub")
|
||||
|
||||
@staticmethod
|
||||
def extract_image_links(chapters):
|
||||
imgs = []
|
||||
for chapter in chapters:
|
||||
chapter_imgs = [urljoin(chapter['asset_base_url'], img_url) for img_url in chapter['images']]
|
||||
imgs.extend(chapter_imgs)
|
||||
return imgs
|
||||
|
||||
|
||||
# MAIN
|
||||
if __name__ == "__main__":
|
||||
|
|
@ -1078,9 +1073,9 @@ if __name__ == "__main__":
|
|||
help="Prevent your session data to be saved into `cookies.json` file."
|
||||
)
|
||||
arguments.add_argument(
|
||||
"--no-kindle", dest="no_kindle", action='store_true',
|
||||
help="Remove some CSS rules that block overflow on `table` and `pre` elements."
|
||||
" Use this option if you're not going to export the EPUB to E-Readers like Amazon Kindle."
|
||||
"--kindle", dest="kindle", action='store_true',
|
||||
help="Add some CSS rules that block overflow on `table` and `pre` elements."
|
||||
" Use this option if you're going to export the EPUB to E-Readers like Amazon Kindle."
|
||||
)
|
||||
arguments.add_argument(
|
||||
"--preserve-log", dest="log", action='store_true', help="Leave the `info_XXXXXXXXXXXXX.log`"
|
||||
|
|
@ -1094,7 +1089,6 @@ if __name__ == "__main__":
|
|||
)
|
||||
|
||||
args_parsed = arguments.parse_args()
|
||||
|
||||
if args_parsed.cred or args_parsed.login:
|
||||
user_email = ""
|
||||
pre_cred = ""
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue