Compare commits
64 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 6612a2f58d | |||
| 1287abeeaf | |||
| eecc0df1ac | |||
| 0183f8f462 | |||
| 674b5f4336 | |||
| 7709e58d64 | |||
| 7695d35449 | |||
| 8ca30f3a3c | |||
| 13a50e2ba2 | |||
| d17dc43ab2 | |||
| de6faa7af1 | |||
| 468512a8cd | |||
| 4edca28c53 | |||
| 2a7a4f5b34 | |||
| 0a3944e54d | |||
| 6b42094db5 | |||
| 937185412a | |||
| 5d20d56e48 | |||
| 9087429501 | |||
| cc905ff2d9 | |||
| eadc54ad25 | |||
| 579bc16be5 | |||
| aae2c6b3d4 | |||
| 705473198f | |||
| b741c0a9e9 | |||
| a6bee88053 | |||
| 1e050e1960 | |||
| 28371817db | |||
| 7ab5db39d0 | |||
| 9a5c4b6865 | |||
| fbe576ffcb | |||
| fcad5067b9 | |||
| 1b8ce1d560 | |||
| 16beb15c43 | |||
| be25e6dbdb | |||
| a13e2f6f1f | |||
| 4b08165328 | |||
| e5b143d9a8 | |||
| 8e5a8e6712 | |||
| 5efbcdcebb | |||
| 189fe58bf2 | |||
| 1575ec1bf0 | |||
| d5d6a5962b | |||
| 420d5aa624 | |||
| a22fa63c4e | |||
| 52b2a595b4 | |||
| afa1ba7c1f | |||
| f725f04223 | |||
| 3afb72b872 | |||
| 6dd9b6ce01 | |||
| fc1b6f6227 | |||
| 7d4c9e53c6 | |||
| f8b6181988 | |||
| dbdbc5f19e | |||
| 44193e0d26 | |||
| a9918a78cf | |||
| 47bb839d7a | |||
| 1b30f8ecf9 | |||
| 8e28a0cac0 | |||
| eb2badbbd0 | |||
| 0d1db4b09e | |||
| 167ee9ac69 | |||
| 83f816f104 | |||
| 9eb15c09dc |
@@ -1,10 +0,0 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
end_of_line = lf
|
||||
insert_final_newline = true
|
||||
|
||||
[*.py]
|
||||
charset = utf-8
|
||||
indent_style = space
|
||||
indent_size = 4
|
||||
@@ -0,0 +1 @@
|
||||
open_collective: camelot
|
||||
@@ -0,0 +1,48 @@
|
||||
---
|
||||
name: Bug report
|
||||
about: Please follow this template to submit bug reports.
|
||||
title: ''
|
||||
labels: bug
|
||||
assignees: ''
|
||||
|
||||
---
|
||||
|
||||
<!-- Please read the filing issues section of the contributor's guide first: https://camelot-py.readthedocs.io/en/master/dev/contributing.html -->
|
||||
|
||||
**Describe the bug**
|
||||
A clear and concise description of what the bug is.
|
||||
|
||||
**Steps to reproduce the bug**
|
||||
Steps used to install `camelot`:
|
||||
1. Add step here (you can add more steps too)
|
||||
|
||||
Steps to reproduce the behavior:
|
||||
1. Add step here (you can add more steps too)
|
||||
|
||||
**Expected behavior**
|
||||
A clear and concise description of what you expected to happen.
|
||||
|
||||
**Code**
|
||||
Add the Camelot code snippet that you used.
|
||||
```
|
||||
import camelot
|
||||
|
||||
# add your code here
|
||||
```
|
||||
|
||||
**PDF**
|
||||
Add the PDF file that you want to extract tables from.
|
||||
|
||||
**Screenshots**
|
||||
If applicable, add screenshots to help explain your problem.
|
||||
|
||||
**Environment**
|
||||
- OS: [e.g. MacOS]
|
||||
- Python version:
|
||||
- Numpy version:
|
||||
- OpenCV version:
|
||||
- Ghostscript version:
|
||||
- Camelot version:
|
||||
|
||||
**Additional context**
|
||||
Add any other context about the problem here.
|
||||
@@ -0,0 +1,27 @@
|
||||
# .readthedocs.yml
|
||||
# Read the Docs configuration file
|
||||
# See https://docs.readthedocs.io/en/stable/config-file/v2.html for details
|
||||
|
||||
# Required
|
||||
version: 2
|
||||
|
||||
# Build documentation in the docs/ directory with Sphinx
|
||||
sphinx:
|
||||
configuration: docs/conf.py
|
||||
|
||||
# Build documentation with MkDocs
|
||||
#mkdocs:
|
||||
# configuration: mkdocs.yml
|
||||
|
||||
# Optionally build your docs in additional formats such as PDF
|
||||
formats:
|
||||
- pdf
|
||||
|
||||
# Optionally set the version of Python and requirements required to build your docs
|
||||
python:
|
||||
version: 3.8
|
||||
install:
|
||||
- method: pip
|
||||
path: .
|
||||
extra_requirements:
|
||||
- dev
|
||||
@@ -8,14 +8,6 @@ install:
|
||||
- make install
|
||||
jobs:
|
||||
include:
|
||||
- stage: test
|
||||
script:
|
||||
- make test
|
||||
python: '2.7'
|
||||
- stage: test
|
||||
script:
|
||||
- make test
|
||||
python: '3.5'
|
||||
- stage: test
|
||||
script:
|
||||
- make test
|
||||
@@ -25,8 +17,13 @@ jobs:
|
||||
- make test
|
||||
python: '3.7'
|
||||
dist: xenial
|
||||
- stage: test
|
||||
script:
|
||||
- make test
|
||||
python: '3.8'
|
||||
dist: xenial
|
||||
- stage: coverage
|
||||
python: '3.6'
|
||||
python: '3.8'
|
||||
script:
|
||||
- make test
|
||||
- codecov --verbose
|
||||
|
||||
@@ -23,7 +23,7 @@ A great way to start contributing to Camelot is to pick an issue tagged with the
|
||||
To install the dependencies needed for development, you can use pip:
|
||||
|
||||
<pre>
|
||||
$ pip install camelot-py[dev]
|
||||
$ pip install "camelot-py[dev]"
|
||||
</pre>
|
||||
|
||||
Alternatively, you can clone the project repository, and install using pip:
|
||||
|
||||
@@ -4,6 +4,30 @@ Release History
|
||||
master
|
||||
------
|
||||
|
||||
0.8.2 (2020-07-27)
|
||||
------------------
|
||||
|
||||
* Revert the changes in `0.8.1`.
|
||||
|
||||
0.8.1 (2020-07-21)
|
||||
------------------
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
* [#169](https://github.com/camelot-dev/camelot/issues/169) Fix import error caused by `pdfminer.six==20200720`. [#171](https://github.com/camelot-dev/camelot/pull/171) by Vinayak Mehta.
|
||||
|
||||
0.8.0 (2020-05-24)
|
||||
------------------
|
||||
|
||||
**Improvements**
|
||||
|
||||
* Drop Python 2 support!
|
||||
* Remove Python 2.7 and 3.5 support.
|
||||
* Replace all instances of `.format` with f-strings.
|
||||
* Remove all `__future__` imports.
|
||||
* Fix HTTP 403 forbidden exception in read_pdf(url) and remove Python 2 urllib support.
|
||||
* Fix test data.
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
* Fix library discovery on Windows. [#32](https://github.com/camelot-dev/camelot/pull/32) by [KOLANICH](https://github.com/KOLANICH).
|
||||
|
||||
@@ -1,12 +1,7 @@
|
||||
MIT License
|
||||
|
||||
Modifications:
|
||||
|
||||
Copyright (c) 2019 Camelot Developers
|
||||
|
||||
Original project:
|
||||
|
||||
Copyright (c) 2018 Peeply Private Ltd (Singapore)
|
||||
Copyright (c) 2019-2020 Camelot Developers
|
||||
Copyright (c) 2018-2019 Peeply Private Ltd (Singapore)
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
|
||||
@@ -7,16 +7,16 @@
|
||||
[](https://travis-ci.org/camelot-dev/camelot) [](https://camelot-py.readthedocs.io/en/master/)
|
||||
[](https://codecov.io/github/camelot-dev/camelot?branch=master)
|
||||
[](https://pypi.org/project/camelot-py/) [](https://pypi.org/project/camelot-py/) [](https://pypi.org/project/camelot-py/) [](https://gitter.im/camelot-dev/Lobby)
|
||||
[](https://github.com/ambv/black)
|
||||
[](https://github.com/ambv/black) [](https://deepsource.io/gh/camelot-dev/camelot/?ref=repository-badge)
|
||||
|
||||
|
||||
**Camelot** is a Python library that makes it easy for *anyone* to extract tables from PDF files!
|
||||
**Camelot** is a Python library that can help you extract tables from PDFs!
|
||||
|
||||
**Note:** You can also check out [Excalibur](https://github.com/camelot-dev/excalibur), which is a web interface for Camelot!
|
||||
**Note:** You can also check out [Excalibur](https://github.com/camelot-dev/excalibur), the web interface to Camelot!
|
||||
|
||||
---
|
||||
|
||||
**Here's how you can extract tables from PDF files.** Check out the PDF used in this example [here](https://github.com/camelot-dev/camelot/blob/master/docs/_static/pdf/foo.pdf).
|
||||
**Here's how you can extract tables from PDFs.** You can check out the PDF used in this example [here](https://github.com/camelot-dev/camelot/blob/master/docs/_static/pdf/foo.pdf).
|
||||
|
||||
<pre>
|
||||
>>> import camelot
|
||||
@@ -46,24 +46,27 @@
|
||||
| 2032_2 | 0.17 | 57.8 | 21.7% | 0.3% | 2.7% | 1.2% |
|
||||
| 4171_1 | 0.07 | 173.9 | 58.1% | 1.6% | 2.1% | 0.5% |
|
||||
|
||||
There's a [command-line interface](https://camelot-py.readthedocs.io/en/master/user/cli.html) too!
|
||||
Camelot also comes packaged with a [command-line interface](https://camelot-py.readthedocs.io/en/master/user/cli.html)!
|
||||
|
||||
**Note:** Camelot only works with text-based PDFs and not scanned documents. (As Tabula [explains](https://github.com/tabulapdf/tabula#why-tabula), "If you can click and drag to select text in your table in a PDF viewer, then your PDF is text-based".)
|
||||
|
||||
## Why Camelot?
|
||||
|
||||
- **You are in control.**: Unlike other libraries and tools which either give a nice output or fail miserably (with no in-between), Camelot gives you the power to tweak table extraction. (This is important since everything in the real world, including PDF table extraction, is fuzzy.)
|
||||
- *Bad* tables can be discarded based on **metrics** like accuracy and whitespace, without ever having to manually look at each table.
|
||||
- Each table is a **pandas DataFrame**, which seamlessly integrates into [ETL and data analysis workflows](https://gist.github.com/vinayak-mehta/e5949f7c2410a0e12f25d3682dc9e873).
|
||||
- **Export** to multiple formats, including JSON, Excel, HTML and Sqlite.
|
||||
- **Configurability**: Camelot gives you control over the table extraction process with its [tweakable settings](https://camelot-py.readthedocs.io/en/master/user/advanced.html).
|
||||
- **Metrics**: Bad tables can be discarded based on metrics like accuracy and whitespace, without having to manually look at each table.
|
||||
- **Output**: Each table is extracted into a **pandas DataFrame**, which seamlessly integrates into [ETL and data analysis workflows](https://gist.github.com/vinayak-mehta/e5949f7c2410a0e12f25d3682dc9e873). You can also export tables to multiple formats, which include CSV, JSON, Excel, HTML and Sqlite.
|
||||
|
||||
See [comparison with other PDF table extraction libraries and tools](https://github.com/camelot-dev/camelot/wiki/Comparison-with-other-PDF-Table-Extraction-libraries-and-tools).
|
||||
See [comparison with similar libraries and tools](https://github.com/camelot-dev/camelot/wiki/Comparison-with-other-PDF-Table-Extraction-libraries-and-tools).
|
||||
|
||||
## Support the development
|
||||
|
||||
If Camelot has helped you, please consider supporting its development with a one-time or monthly donation [on OpenCollective](https://opencollective.com/camelot).
|
||||
|
||||
## Installation
|
||||
|
||||
### Using conda
|
||||
|
||||
The easiest way to install Camelot is to install it with [conda](https://conda.io/docs/), which is a package manager and environment management system for the [Anaconda](http://docs.continuum.io/anaconda/) distribution.
|
||||
The easiest way to install Camelot is with [conda](https://conda.io/docs/), which is a package manager and environment management system for the [Anaconda](http://docs.continuum.io/anaconda/) distribution.
|
||||
|
||||
<pre>
|
||||
$ conda install -c conda-forge camelot-py
|
||||
@@ -71,10 +74,10 @@ $ conda install -c conda-forge camelot-py
|
||||
|
||||
### Using pip
|
||||
|
||||
After [installing the dependencies](https://camelot-py.readthedocs.io/en/master/user/install-deps.html) ([tk](https://packages.ubuntu.com/bionic/python/python-tk) and [ghostscript](https://www.ghostscript.com/)), you can simply use pip to install Camelot:
|
||||
After [installing the dependencies](https://camelot-py.readthedocs.io/en/master/user/install-deps.html) ([tk](https://packages.ubuntu.com/bionic/python/python-tk) and [ghostscript](https://www.ghostscript.com/)), you can also just use pip to install Camelot:
|
||||
|
||||
<pre>
|
||||
$ pip install camelot-py[cv]
|
||||
$ pip install "camelot-py[cv]"
|
||||
</pre>
|
||||
|
||||
### From the source code
|
||||
@@ -94,35 +97,15 @@ $ pip install ".[cv]"
|
||||
|
||||
## Documentation
|
||||
|
||||
Great documentation is available at [http://camelot-py.readthedocs.io/](http://camelot-py.readthedocs.io/).
|
||||
The documentation is available at [http://camelot-py.readthedocs.io/](http://camelot-py.readthedocs.io/).
|
||||
|
||||
## Development
|
||||
## Wrappers
|
||||
|
||||
The [Contributor's Guide](https://camelot-py.readthedocs.io/en/master/dev/contributing.html) has detailed information about contributing code, documentation, tests and more. We've included some basic information in this README.
|
||||
- [camelot-php](https://github.com/randomstate/camelot-php) provides a [PHP](https://www.php.net/) wrapper on Camelot.
|
||||
|
||||
### Source code
|
||||
## Contributing
|
||||
|
||||
You can check the latest sources with:
|
||||
|
||||
<pre>
|
||||
$ git clone https://www.github.com/camelot-dev/camelot
|
||||
</pre>
|
||||
|
||||
### Setting up a development environment
|
||||
|
||||
You can install the development dependencies easily, using pip:
|
||||
|
||||
<pre>
|
||||
$ pip install camelot-py[dev]
|
||||
</pre>
|
||||
|
||||
### Testing
|
||||
|
||||
After installation, you can run tests using:
|
||||
|
||||
<pre>
|
||||
$ python setup.py test
|
||||
</pre>
|
||||
The [Contributor's Guide](https://camelot-py.readthedocs.io/en/master/dev/contributing.html) has detailed information about contributing issues, documentation, code, and tests.
|
||||
|
||||
## Versioning
|
||||
|
||||
@@ -131,9 +114,3 @@ Camelot uses [Semantic Versioning](https://semver.org/). For the available versi
|
||||
## License
|
||||
|
||||
This project is licensed under the MIT License, see the [LICENSE](https://github.com/camelot-dev/camelot/blob/master/LICENSE) file for details.
|
||||
|
||||
## Support the development
|
||||
|
||||
You can support our work on Camelot with a one-time or monthly donation [on OpenCollective](https://opencollective.com/camelot). Organizations who use camelot can also sponsor the project for an acknowledgement on [our documentation site](https://camelot-py.readthedocs.io/en/master/) and this README.
|
||||
|
||||
Special thanks to all the users, organizations and contributors that support Camelot!
|
||||
|
||||
@@ -1,7 +1,5 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from __future__ import absolute_import
|
||||
|
||||
|
||||
__all__ = ("main",)
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
VERSION = (0, 7, 3)
|
||||
VERSION = (0, 8, 2)
|
||||
PRERELEASE = None # alpha, beta or rc
|
||||
REVISION = None
|
||||
|
||||
@@ -8,9 +8,9 @@ REVISION = None
|
||||
def generate_version(version, prerelease=None, revision=None):
|
||||
version_parts = [".".join(map(str, version))]
|
||||
if prerelease is not None:
|
||||
version_parts.append("-{}".format(prerelease))
|
||||
version_parts.append(f"-{prerelease}")
|
||||
if revision is not None:
|
||||
version_parts.append(".{}".format(revision))
|
||||
version_parts.append(f".{revision}")
|
||||
return "".join(version_parts)
|
||||
|
||||
|
||||
|
||||
@@ -204,7 +204,7 @@ def lattice(c, *args, **kwargs):
|
||||
tables = read_pdf(
|
||||
filepath, pages=pages, flavor="lattice", suppress_stdout=quiet, **kwargs
|
||||
)
|
||||
click.echo("Found {} tables".format(tables.n))
|
||||
click.echo(f"Found {tables.n} tables")
|
||||
if plot_type is not None:
|
||||
for table in tables:
|
||||
plot(table, kind=plot_type)
|
||||
@@ -295,7 +295,7 @@ def stream(c, *args, **kwargs):
|
||||
tables = read_pdf(
|
||||
filepath, pages=pages, flavor="stream", suppress_stdout=quiet, **kwargs
|
||||
)
|
||||
click.echo("Found {} tables".format(tables.n))
|
||||
click.echo(f"Found {tables.n} tables")
|
||||
if plot_type is not None:
|
||||
for table in tables:
|
||||
plot(table, kind=plot_type)
|
||||
|
||||
@@ -52,13 +52,10 @@ class TextEdge(object):
|
||||
self.is_valid = False
|
||||
|
||||
def __repr__(self):
|
||||
return "<TextEdge x={} y0={} y1={} align={} valid={}>".format(
|
||||
round(self.x, 2),
|
||||
round(self.y0, 2),
|
||||
round(self.y1, 2),
|
||||
self.align,
|
||||
self.is_valid,
|
||||
)
|
||||
x = round(self.x, 2)
|
||||
y0 = round(self.y0, 2)
|
||||
y1 = round(self.y1, 2)
|
||||
return f"<TextEdge x={x} y0={y0} y1={y1} align={self.align} valid={self.is_valid}>"
|
||||
|
||||
def update_coords(self, x, y0, edge_tol=50):
|
||||
"""Updates the text edge's x and bottom y coordinates and sets
|
||||
@@ -291,9 +288,11 @@ class Cell(object):
|
||||
self._text = ""
|
||||
|
||||
def __repr__(self):
|
||||
return "<Cell x1={} y1={} x2={} y2={}>".format(
|
||||
round(self.x1, 2), round(self.y1, 2), round(self.x2, 2), round(self.y2, 2)
|
||||
)
|
||||
x1 = round(self.x1, 2)
|
||||
y1 = round(self.y1, 2)
|
||||
x2 = round(self.x2, 2)
|
||||
y2 = round(self.y2, 2)
|
||||
return f"<Cell x1={x1} y1={y1} x2={x2} y2={y2}>"
|
||||
|
||||
@property
|
||||
def text(self):
|
||||
@@ -351,7 +350,7 @@ class Table(object):
|
||||
self.page = None
|
||||
|
||||
def __repr__(self):
|
||||
return "<{} shape={}>".format(self.__class__.__name__, self.shape)
|
||||
return f"<{self.__class__.__name__} shape={self.shape}>"
|
||||
|
||||
def __lt__(self, other):
|
||||
if self.page == other.page:
|
||||
@@ -612,7 +611,7 @@ class Table(object):
|
||||
|
||||
"""
|
||||
kw = {
|
||||
"sheet_name": "page-{}-table-{}".format(self.page, self.order),
|
||||
"sheet_name": f"page-{self.page}-table-{self.order}",
|
||||
"encoding": "utf-8",
|
||||
}
|
||||
kw.update(kwargs)
|
||||
@@ -632,7 +631,7 @@ class Table(object):
|
||||
|
||||
"""
|
||||
html_string = self.df.to_html(**kwargs)
|
||||
with open(path, "w") as f:
|
||||
with open(path, "w", encoding="utf-8") as f:
|
||||
f.write(html_string)
|
||||
|
||||
def to_sqlite(self, path, **kwargs):
|
||||
@@ -649,7 +648,7 @@ class Table(object):
|
||||
kw = {"if_exists": "replace", "index": False}
|
||||
kw.update(kwargs)
|
||||
conn = sqlite3.connect(path)
|
||||
table_name = "page-{}-table-{}".format(self.page, self.order)
|
||||
table_name = f"page-{self.page}-table-{self.order}"
|
||||
self.df.to_sql(table_name, conn, **kw)
|
||||
conn.commit()
|
||||
conn.close()
|
||||
@@ -670,7 +669,7 @@ class TableList(object):
|
||||
self._tables = tables
|
||||
|
||||
def __repr__(self):
|
||||
return "<{} n={}>".format(self.__class__.__name__, self.n)
|
||||
return f"<{self.__class__.__name__} n={self.n}>"
|
||||
|
||||
def __len__(self):
|
||||
return len(self._tables)
|
||||
@@ -680,7 +679,7 @@ class TableList(object):
|
||||
|
||||
@staticmethod
|
||||
def _format_func(table, f):
|
||||
return getattr(table, "to_{}".format(f))
|
||||
return getattr(table, f"to_{f}")
|
||||
|
||||
@property
|
||||
def n(self):
|
||||
@@ -691,9 +690,7 @@ class TableList(object):
|
||||
root = kwargs.get("root")
|
||||
ext = kwargs.get("ext")
|
||||
for table in self._tables:
|
||||
filename = os.path.join(
|
||||
"{}-page-{}-table-{}{}".format(root, table.page, table.order, ext)
|
||||
)
|
||||
filename = f"{root}-page-{table.page}-table-{table.order}{ext}"
|
||||
filepath = os.path.join(dirname, filename)
|
||||
to_format = self._format_func(table, f)
|
||||
to_format(filepath)
|
||||
@@ -706,9 +703,7 @@ class TableList(object):
|
||||
zipname = os.path.join(os.path.dirname(path), root) + ".zip"
|
||||
with zipfile.ZipFile(zipname, "w", allowZip64=True) as z:
|
||||
for table in self._tables:
|
||||
filename = os.path.join(
|
||||
"{}-page-{}-table-{}{}".format(root, table.page, table.order, ext)
|
||||
)
|
||||
filename = f"{root}-page-{table.page}-table-{table.order}{ext}"
|
||||
filepath = os.path.join(dirname, filename)
|
||||
z.write(filepath, os.path.basename(filepath))
|
||||
|
||||
@@ -741,7 +736,7 @@ class TableList(object):
|
||||
filepath = os.path.join(dirname, basename)
|
||||
writer = pd.ExcelWriter(filepath)
|
||||
for table in self._tables:
|
||||
sheet_name = "page-{}-table-{}".format(table.page, table.order)
|
||||
sheet_name = f"page-{table.page}-table-{table.order}"
|
||||
table.df.to_excel(writer, sheet_name=sheet_name, encoding="utf-8")
|
||||
writer.save()
|
||||
if compress:
|
||||
|
||||
@@ -81,6 +81,7 @@ def delete_instance(instance):
|
||||
"""
|
||||
return libgs.gsapi_delete_instance(instance)
|
||||
|
||||
|
||||
if sys.platform == "win32":
|
||||
c_stdstream_call_t = WINFUNCTYPE(c_int, gs_main_instance, POINTER(c_char), c_int)
|
||||
else:
|
||||
@@ -247,7 +248,10 @@ if sys.platform == "win32":
|
||||
libgs = __win32_finddll()
|
||||
if not libgs:
|
||||
import ctypes.util
|
||||
libgs = ctypes.util.find_library("".join(("gsdll", str(ctypes.sizeof(ctypes.c_voidp) * 8), ".dll"))) # finds in %PATH%
|
||||
|
||||
libgs = ctypes.util.find_library(
|
||||
"".join(("gsdll", str(ctypes.sizeof(ctypes.c_voidp) * 8), ".dll"))
|
||||
) # finds in %PATH%
|
||||
if not libgs:
|
||||
raise RuntimeError("Please make sure that Ghostscript is installed")
|
||||
libgs = windll.LoadLibrary(libgs)
|
||||
|
||||
@@ -6,7 +6,7 @@ import sys
|
||||
from PyPDF2 import PdfFileReader, PdfFileWriter
|
||||
|
||||
from .core import TableList
|
||||
from .parsers import Stream, Lattice
|
||||
from .parsers import Lattice, Stream, LatticeOCR, StreamOCR
|
||||
from .utils import (
|
||||
TemporaryDirectory,
|
||||
get_page_layout,
|
||||
@@ -70,7 +70,8 @@ class PDFHandler(object):
|
||||
if pages == "1":
|
||||
page_numbers.append({"start": 1, "end": 1})
|
||||
else:
|
||||
infile = PdfFileReader(open(filepath, "rb"), strict=False)
|
||||
instream = open(filepath, "rb")
|
||||
infile = PdfFileReader(instream, strict=False)
|
||||
if infile.isEncrypted:
|
||||
infile.decrypt(self.password)
|
||||
if pages == "all":
|
||||
@@ -84,6 +85,7 @@ class PDFHandler(object):
|
||||
page_numbers.append({"start": int(a), "end": int(b)})
|
||||
else:
|
||||
page_numbers.append({"start": int(r), "end": int(r)})
|
||||
instream.close()
|
||||
P = []
|
||||
for p in page_numbers:
|
||||
P.extend(range(p["start"], p["end"] + 1))
|
||||
@@ -106,7 +108,7 @@ class PDFHandler(object):
|
||||
infile = PdfFileReader(fileobj, strict=False)
|
||||
if infile.isEncrypted:
|
||||
infile.decrypt(self.password)
|
||||
fpath = os.path.join(temp, "page-{0}.pdf".format(page))
|
||||
fpath = os.path.join(temp, f"page-{page}.pdf")
|
||||
froot, fext = os.path.splitext(fpath)
|
||||
p = infile.getPage(page - 1)
|
||||
outfile = PdfFileWriter()
|
||||
@@ -122,7 +124,8 @@ class PDFHandler(object):
|
||||
if rotation != "":
|
||||
fpath_new = "".join([froot.replace("page", "p"), "_rotated", fext])
|
||||
os.rename(fpath, fpath_new)
|
||||
infile = PdfFileReader(open(fpath_new, "rb"), strict=False)
|
||||
instream = open(fpath_new, "rb")
|
||||
infile = PdfFileReader(instream, strict=False)
|
||||
if infile.isEncrypted:
|
||||
infile.decrypt(self.password)
|
||||
outfile = PdfFileWriter()
|
||||
@@ -134,6 +137,7 @@ class PDFHandler(object):
|
||||
outfile.addPage(p)
|
||||
with open(fpath, "wb") as f:
|
||||
outfile.write(f)
|
||||
instream.close()
|
||||
|
||||
def parse(
|
||||
self, flavor="lattice", suppress_stdout=False, layout_kwargs={}, **kwargs
|
||||
@@ -159,14 +163,19 @@ class PDFHandler(object):
|
||||
List of tables found in PDF.
|
||||
|
||||
"""
|
||||
parsers = {
|
||||
"lattice": Lattice,
|
||||
"stream": Stream,
|
||||
"lattice_ocr": LatticeOCR,
|
||||
"stream_ocr": StreamOCR,
|
||||
}
|
||||
|
||||
tables = []
|
||||
with TemporaryDirectory() as tempdir:
|
||||
for p in self.pages:
|
||||
self._save_page(self.filepath, p, tempdir)
|
||||
pages = [
|
||||
os.path.join(tempdir, "page-{0}.pdf".format(p)) for p in self.pages
|
||||
]
|
||||
parser = Lattice(**kwargs) if flavor == "lattice" else Stream(**kwargs)
|
||||
pages = [os.path.join(tempdir, f"page-{p}.pdf") for p in self.pages]
|
||||
parser = parsers[flavor](**kwargs)
|
||||
for p in pages:
|
||||
t = parser.extract_tables(
|
||||
p, suppress_stdout=suppress_stdout, layout_kwargs=layout_kwargs
|
||||
|
||||
@@ -1,7 +1,5 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from __future__ import division
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
|
||||
@@ -98,9 +98,10 @@ def read_pdf(
|
||||
tables : camelot.core.TableList
|
||||
|
||||
"""
|
||||
if flavor not in ["lattice", "stream"]:
|
||||
if flavor not in ["lattice", "stream", "lattice_ocr", "stream_ocr"]:
|
||||
raise NotImplementedError(
|
||||
"Unknown flavor specified." " Use either 'lattice' or 'stream'"
|
||||
"Unknown flavor specified. Use one of the following: "
|
||||
"'lattice', 'stream', 'lattice_ocr', 'stream_ocr'"
|
||||
)
|
||||
|
||||
with warnings.catch_warnings():
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from .stream import Stream
|
||||
from .lattice import Lattice
|
||||
from .stream import Stream
|
||||
from .lattice_ocr import LatticeOCR
|
||||
from .stream_ocr import StreamOCR
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from __future__ import division
|
||||
import os
|
||||
import sys
|
||||
import copy
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
import os
|
||||
import copy
|
||||
import logging
|
||||
import subprocess
|
||||
|
||||
try:
|
||||
import easyocr
|
||||
except ImportError:
|
||||
_HAS_EASYOCR = False
|
||||
else:
|
||||
_HAS_EASYOCR = True
|
||||
|
||||
import pandas as pd
|
||||
from PIL import Image
|
||||
|
||||
from .base import BaseParser
|
||||
from ..core import Table
|
||||
from ..utils import TemporaryDirectory, merge_close_lines, scale_pdf, segments_in_bbox
|
||||
from ..image_processing import (
|
||||
adaptive_threshold,
|
||||
find_lines,
|
||||
find_contours,
|
||||
find_joints,
|
||||
)
|
||||
|
||||
|
||||
logger = logging.getLogger("camelot")
|
||||
|
||||
|
||||
class LatticeOCR(BaseParser):
|
||||
def __init__(
|
||||
self,
|
||||
table_regions=None,
|
||||
table_areas=None,
|
||||
line_scale=15,
|
||||
line_tol=2,
|
||||
joint_tol=2,
|
||||
threshold_blocksize=15,
|
||||
threshold_constant=-2,
|
||||
iterations=0,
|
||||
resolution=300,
|
||||
):
|
||||
self.table_regions = table_regions
|
||||
self.table_areas = table_areas
|
||||
self.line_scale = line_scale
|
||||
self.line_tol = line_tol
|
||||
self.joint_tol = joint_tol
|
||||
self.threshold_blocksize = threshold_blocksize
|
||||
self.threshold_constant = threshold_constant
|
||||
self.iterations = iterations
|
||||
self.resolution = resolution
|
||||
|
||||
if _HAS_EASYOCR:
|
||||
self.reader = easyocr.Reader(['en'], gpu=False)
|
||||
else:
|
||||
raise ImportError("easyocr is required to run OCR on image-based PDFs.")
|
||||
|
||||
def _generate_image(self):
|
||||
from ..ext.ghostscript import Ghostscript
|
||||
|
||||
self.imagename = "".join([self.rootname, ".png"])
|
||||
gs_call = "-q -sDEVICE=png16m -o {} -r900 {}".format(
|
||||
self.imagename, self.filename
|
||||
)
|
||||
gs_call = gs_call.encode().split()
|
||||
null = open(os.devnull, "wb")
|
||||
with Ghostscript(*gs_call, stdout=null) as gs:
|
||||
pass
|
||||
null.close()
|
||||
|
||||
def _generate_table_bbox(self):
|
||||
def scale_areas(areas, scalers):
|
||||
scaled_areas = []
|
||||
for area in areas:
|
||||
x1, y1, x2, y2 = area.split(",")
|
||||
x1 = float(x1)
|
||||
y1 = float(y1)
|
||||
x2 = float(x2)
|
||||
y2 = float(y2)
|
||||
x1, y1, x2, y2 = scale_pdf((x1, y1, x2, y2), scalers)
|
||||
scaled_areas.append((x1, y1, abs(x2 - x1), abs(y2 - y1)))
|
||||
return scaled_areas
|
||||
|
||||
self.image, self.threshold = adaptive_threshold(
|
||||
self.imagename, blocksize=self.threshold_blocksize, c=self.threshold_constant
|
||||
)
|
||||
|
||||
image_width = self.image.shape[1]
|
||||
image_height = self.image.shape[0]
|
||||
image_width_scaler = image_width / float(self.pdf_width)
|
||||
image_height_scaler = image_height / float(self.pdf_height)
|
||||
image_scalers = (image_width_scaler, image_height_scaler, self.pdf_height)
|
||||
|
||||
if self.table_areas is None:
|
||||
regions = None
|
||||
if self.table_regions is not None:
|
||||
regions = scale_areas(self.table_regions, image_scalers)
|
||||
|
||||
vertical_mask, vertical_segments = find_lines(
|
||||
self.threshold,
|
||||
regions=regions,
|
||||
direction="vertical",
|
||||
line_scale=self.line_scale,
|
||||
iterations=self.iterations,
|
||||
)
|
||||
horizontal_mask, horizontal_segments = find_lines(
|
||||
self.threshold,
|
||||
regions=regions,
|
||||
direction="horizontal",
|
||||
line_scale=self.line_scale,
|
||||
iterations=self.iterations,
|
||||
)
|
||||
|
||||
contours = find_contours(vertical_mask, horizontal_mask)
|
||||
table_bbox = find_joints(contours, vertical_mask, horizontal_mask)
|
||||
else:
|
||||
vertical_mask, vertical_segments = find_lines(
|
||||
self.threshold,
|
||||
direction="vertical",
|
||||
line_scale=self.line_scale,
|
||||
iterations=self.iterations,
|
||||
)
|
||||
horizontal_mask, horizontal_segments = find_lines(
|
||||
self.threshold,
|
||||
direction="horizontal",
|
||||
line_scale=self.line_scale,
|
||||
iterations=self.iterations,
|
||||
)
|
||||
|
||||
areas = scale_areas(self.table_areas, image_scalers)
|
||||
table_bbox = find_joints(areas, vertical_mask, horizontal_mask)
|
||||
|
||||
self.table_bbox_unscaled = copy.deepcopy(table_bbox)
|
||||
|
||||
self.table_bbox = table_bbox
|
||||
self.vertical_segments = vertical_segments
|
||||
self.horizontal_segments = horizontal_segments
|
||||
|
||||
def _generate_columns_and_rows(self, table_idx, tk):
|
||||
cols, rows = zip(*self.table_bbox[tk])
|
||||
cols, rows = list(cols), list(rows)
|
||||
cols.extend([tk[0], tk[2]])
|
||||
rows.extend([tk[1], tk[3]])
|
||||
# sort horizontal and vertical segments
|
||||
cols = merge_close_lines(sorted(cols), line_tol=self.line_tol)
|
||||
rows = merge_close_lines(sorted(rows), line_tol=self.line_tol)
|
||||
# make grid using x and y coord of shortlisted rows and cols
|
||||
cols = [(cols[i], cols[i + 1]) for i in range(0, len(cols) - 1)]
|
||||
rows = [(rows[i], rows[i + 1]) for i in range(0, len(rows) - 1)]
|
||||
|
||||
return cols, rows
|
||||
|
||||
|
||||
def _generate_table(self, table_idx, cols, rows, **kwargs):
|
||||
table = Table(cols, rows)
|
||||
# set table edges to True using ver+hor lines
|
||||
table = table.set_edges(self.vertical_segments, self.horizontal_segments, joint_tol=self.joint_tol)
|
||||
# set table border edges to True
|
||||
table = table.set_border()
|
||||
# set spanning cells to True
|
||||
table = table.set_span()
|
||||
|
||||
_seen = set()
|
||||
for r_idx in range(len(table.cells)):
|
||||
for c_idx in range(len(table.cells[r_idx])):
|
||||
if (r_idx, c_idx) in _seen:
|
||||
continue
|
||||
|
||||
_seen.add((r_idx, c_idx))
|
||||
|
||||
_r_idx = r_idx
|
||||
_c_idx = c_idx
|
||||
|
||||
if table.cells[r_idx][_c_idx].hspan:
|
||||
while not table.cells[r_idx][_c_idx].right:
|
||||
_c_idx += 1
|
||||
_seen.add((r_idx, _c_idx))
|
||||
|
||||
if table.cells[_r_idx][c_idx].vspan:
|
||||
while not table.cells[_r_idx][c_idx].bottom:
|
||||
_r_idx += 1
|
||||
_seen.add((_r_idx, c_idx))
|
||||
|
||||
for i in range(r_idx, _r_idx + 1):
|
||||
for j in range(c_idx, _c_idx + 1):
|
||||
_seen.add((i, j))
|
||||
|
||||
x1 = int(table.cells[r_idx][c_idx].x1)
|
||||
y1 = int(table.cells[_r_idx][_c_idx].y1)
|
||||
|
||||
x2 = int(table.cells[_r_idx][_c_idx].x2)
|
||||
y2 = int(table.cells[r_idx][c_idx].y2)
|
||||
|
||||
with TemporaryDirectory() as tempdir:
|
||||
temp_image_path = os.path.join(tempdir, f"{table_idx}_{r_idx}_{c_idx}.png")
|
||||
|
||||
cell_image = Image.fromarray(self.image[y2:y1, x1:x2])
|
||||
cell_image.save(temp_image_path)
|
||||
|
||||
text = self.reader.readtext(temp_image_path, detail=0)
|
||||
text = " ".join(text)
|
||||
|
||||
table.cells[r_idx][c_idx].text = text
|
||||
|
||||
data = table.data
|
||||
table.df = pd.DataFrame(data)
|
||||
table.shape = table.df.shape
|
||||
|
||||
table.flavor = "lattice_ocr"
|
||||
table.accuracy = 0
|
||||
table.whitespace = 0
|
||||
table.order = table_idx + 1
|
||||
table.page = int(os.path.basename(self.rootname).replace("page-", ""))
|
||||
|
||||
# for plotting
|
||||
table._text = None
|
||||
table._image = (self.image, self.table_bbox_unscaled)
|
||||
table._segments = (self.vertical_segments, self.horizontal_segments)
|
||||
table._textedges = None
|
||||
|
||||
return table
|
||||
|
||||
def extract_tables(self, filename, suppress_stdout=False, layout_kwargs={}):
|
||||
self._generate_layout(filename, layout_kwargs)
|
||||
if not suppress_stdout:
|
||||
logger.info("Processing {}".format(os.path.basename(self.rootname)))
|
||||
|
||||
self._generate_image()
|
||||
self._generate_table_bbox()
|
||||
|
||||
_tables = []
|
||||
# sort tables based on y-coord
|
||||
for table_idx, tk in enumerate(
|
||||
sorted(self.table_bbox.keys(), key=lambda x: x[1], reverse=True)
|
||||
):
|
||||
cols, rows = self._generate_columns_and_rows(table_idx, tk)
|
||||
table = self._generate_table(table_idx, cols, rows)
|
||||
table._bbox = tk
|
||||
_tables.append(table)
|
||||
|
||||
return _tables
|
||||
@@ -1,6 +1,5 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from __future__ import division
|
||||
import os
|
||||
import logging
|
||||
import warnings
|
||||
@@ -122,6 +121,7 @@ class Stream(BaseParser):
|
||||
row_y = 0
|
||||
rows = []
|
||||
temp = []
|
||||
|
||||
for t in text:
|
||||
# is checking for upright necessary?
|
||||
# if t.get_text().strip() and all([obj.upright for obj in t._objs if
|
||||
@@ -132,8 +132,10 @@ class Stream(BaseParser):
|
||||
temp = []
|
||||
row_y = t.y0
|
||||
temp.append(t)
|
||||
|
||||
rows.append(sorted(temp, key=lambda t: t.x0))
|
||||
__ = rows.pop(0) # TODO: hacky
|
||||
if len(rows) > 1:
|
||||
__ = rows.pop(0) # TODO: hacky
|
||||
return rows
|
||||
|
||||
@staticmethod
|
||||
@@ -346,43 +348,46 @@ class Stream(BaseParser):
|
||||
else:
|
||||
# calculate mode of the list of number of elements in
|
||||
# each row to guess the number of columns
|
||||
ncols = max(set(elements), key=elements.count)
|
||||
if ncols == 1:
|
||||
# if mode is 1, the page usually contains not tables
|
||||
# but there can be cases where the list can be skewed,
|
||||
# try to remove all 1s from list in this case and
|
||||
# see if the list contains elements, if yes, then use
|
||||
# the mode after removing 1s
|
||||
elements = list(filter(lambda x: x != 1, elements))
|
||||
if len(elements):
|
||||
ncols = max(set(elements), key=elements.count)
|
||||
else:
|
||||
warnings.warn(
|
||||
"No tables found in table area {}".format(table_idx + 1)
|
||||
if not len(elements):
|
||||
cols = [(text_x_min, text_x_max)]
|
||||
else:
|
||||
ncols = max(set(elements), key=elements.count)
|
||||
if ncols == 1:
|
||||
# if mode is 1, the page usually contains not tables
|
||||
# but there can be cases where the list can be skewed,
|
||||
# try to remove all 1s from list in this case and
|
||||
# see if the list contains elements, if yes, then use
|
||||
# the mode after removing 1s
|
||||
elements = list(filter(lambda x: x != 1, elements))
|
||||
if len(elements):
|
||||
ncols = max(set(elements), key=elements.count)
|
||||
else:
|
||||
warnings.warn(
|
||||
f"No tables found in table area {table_idx + 1}"
|
||||
)
|
||||
cols = [(t.x0, t.x1) for r in rows_grouped if len(r) == ncols for t in r]
|
||||
cols = self._merge_columns(sorted(cols), column_tol=self.column_tol)
|
||||
inner_text = []
|
||||
for i in range(1, len(cols)):
|
||||
left = cols[i - 1][1]
|
||||
right = cols[i][0]
|
||||
inner_text.extend(
|
||||
[
|
||||
t
|
||||
for direction in self.t_bbox
|
||||
for t in self.t_bbox[direction]
|
||||
if t.x0 > left and t.x1 < right
|
||||
]
|
||||
)
|
||||
cols = [(t.x0, t.x1) for r in rows_grouped if len(r) == ncols for t in r]
|
||||
cols = self._merge_columns(sorted(cols), column_tol=self.column_tol)
|
||||
inner_text = []
|
||||
for i in range(1, len(cols)):
|
||||
left = cols[i - 1][1]
|
||||
right = cols[i][0]
|
||||
inner_text.extend(
|
||||
[
|
||||
t
|
||||
for direction in self.t_bbox
|
||||
for t in self.t_bbox[direction]
|
||||
if t.x0 > left and t.x1 < right
|
||||
]
|
||||
)
|
||||
outer_text = [
|
||||
t
|
||||
for direction in self.t_bbox
|
||||
for t in self.t_bbox[direction]
|
||||
if t.x0 > cols[-1][1] or t.x1 < cols[0][0]
|
||||
]
|
||||
inner_text.extend(outer_text)
|
||||
cols = self._add_columns(cols, inner_text, self.row_tol)
|
||||
cols = self._join_columns(cols, text_x_min, text_x_max)
|
||||
outer_text = [
|
||||
t
|
||||
for direction in self.t_bbox
|
||||
for t in self.t_bbox[direction]
|
||||
if t.x0 > cols[-1][1] or t.x1 < cols[0][0]
|
||||
]
|
||||
inner_text.extend(outer_text)
|
||||
cols = self._add_columns(cols, inner_text, self.row_tol)
|
||||
cols = self._join_columns(cols, text_x_min, text_x_max)
|
||||
|
||||
return cols, rows
|
||||
|
||||
@@ -433,19 +438,19 @@ class Stream(BaseParser):
|
||||
|
||||
def extract_tables(self, filename, suppress_stdout=False, layout_kwargs={}):
|
||||
self._generate_layout(filename, layout_kwargs)
|
||||
base_filename = os.path.basename(self.rootname)
|
||||
|
||||
if not suppress_stdout:
|
||||
logger.info("Processing {}".format(os.path.basename(self.rootname)))
|
||||
logger.info(f"Processing {base_filename}")
|
||||
|
||||
if not self.horizontal_text:
|
||||
if self.images:
|
||||
warnings.warn(
|
||||
"{} is image-based, camelot only works on"
|
||||
" text-based pages.".format(os.path.basename(self.rootname))
|
||||
f"{base_filename} is image-based, camelot only works on"
|
||||
" text-based pages."
|
||||
)
|
||||
else:
|
||||
warnings.warn(
|
||||
"No tables found on {}".format(os.path.basename(self.rootname))
|
||||
)
|
||||
warnings.warn(f"No tables found on {base_filename}")
|
||||
return []
|
||||
|
||||
self._generate_table_bbox()
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from .base import BaseParser
|
||||
|
||||
|
||||
class StreamOCR(BaseParser):
|
||||
pass
|
||||
@@ -35,15 +35,21 @@ class PlotMethods(object):
|
||||
|
||||
if table.flavor == "lattice" and kind in ["textedge"]:
|
||||
raise NotImplementedError(
|
||||
"Lattice flavor does not support kind='{}'".format(kind)
|
||||
f"Lattice flavor does not support kind='{kind}'"
|
||||
)
|
||||
elif table.flavor == "stream" and kind in ["joint", "line"]:
|
||||
raise NotImplementedError(
|
||||
"Stream flavor does not support kind='{}'".format(kind)
|
||||
f"Stream flavor does not support kind='{kind}'"
|
||||
)
|
||||
|
||||
plot_method = getattr(self, kind)
|
||||
return plot_method(table)
|
||||
fig = plot_method(table)
|
||||
|
||||
if filename is not None:
|
||||
fig.savefig(filename)
|
||||
return None
|
||||
|
||||
return fig
|
||||
|
||||
def text(self, table):
|
||||
"""Generates a plot for all text elements present
|
||||
|
||||
@@ -1,9 +1,7 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
from __future__ import division
|
||||
|
||||
import re
|
||||
import os
|
||||
import sys
|
||||
import re
|
||||
import random
|
||||
import shutil
|
||||
import string
|
||||
@@ -29,16 +27,9 @@ from pdfminer.layout import (
|
||||
LTImage,
|
||||
)
|
||||
|
||||
|
||||
PY3 = sys.version_info[0] >= 3
|
||||
if PY3:
|
||||
from urllib.request import urlopen
|
||||
from urllib.parse import urlparse as parse_url
|
||||
from urllib.parse import uses_relative, uses_netloc, uses_params
|
||||
else:
|
||||
from urllib2 import urlopen
|
||||
from urlparse import urlparse as parse_url
|
||||
from urlparse import uses_relative, uses_netloc, uses_params
|
||||
from urllib.request import Request, urlopen
|
||||
from urllib.parse import urlparse as parse_url
|
||||
from urllib.parse import uses_relative, uses_netloc, uses_params
|
||||
|
||||
|
||||
_VALID_URLS = set(uses_relative + uses_netloc + uses_params)
|
||||
@@ -88,13 +79,12 @@ def download_url(url):
|
||||
Temporary filepath.
|
||||
|
||||
"""
|
||||
filename = "{}.pdf".format(random_string(6))
|
||||
filename = f"{random_string(6)}.pdf"
|
||||
with tempfile.NamedTemporaryFile("wb", delete=False) as f:
|
||||
obj = urlopen(url)
|
||||
if PY3:
|
||||
content_type = obj.info().get_content_type()
|
||||
else:
|
||||
content_type = obj.info().getheader("Content-Type")
|
||||
headers = {"User-Agent": "Mozilla/5.0"}
|
||||
request = Request(url, None, headers)
|
||||
obj = urlopen(request)
|
||||
content_type = obj.info().get_content_type()
|
||||
if content_type != "application/pdf":
|
||||
raise NotImplementedError("File format not supported")
|
||||
f.write(obj.read())
|
||||
@@ -103,7 +93,6 @@ def download_url(url):
|
||||
return filepath
|
||||
|
||||
|
||||
stream_kwargs = ["columns", "edge_tol", "row_tol", "column_tol"]
|
||||
lattice_kwargs = [
|
||||
"process_background",
|
||||
"line_scale",
|
||||
@@ -116,6 +105,7 @@ lattice_kwargs = [
|
||||
"iterations",
|
||||
"resolution",
|
||||
]
|
||||
stream_kwargs = ["columns", "edge_tol", "row_tol", "column_tol"]
|
||||
|
||||
|
||||
def validate_input(kwargs, flavor="lattice"):
|
||||
@@ -123,19 +113,17 @@ def validate_input(kwargs, flavor="lattice"):
|
||||
isec = set(parser_kwargs).intersection(set(input_kwargs.keys()))
|
||||
if isec:
|
||||
raise ValueError(
|
||||
"{} cannot be used with flavor='{}'".format(
|
||||
",".join(sorted(isec)), flavor
|
||||
)
|
||||
f"{','.join(sorted(isec))} cannot be used with flavor='{flavor}'"
|
||||
)
|
||||
|
||||
if flavor == "lattice":
|
||||
if flavor in ["lattice", "lattice_ocr"]:
|
||||
check_intersection(stream_kwargs, kwargs)
|
||||
else:
|
||||
check_intersection(lattice_kwargs, kwargs)
|
||||
|
||||
|
||||
def remove_extra(kwargs, flavor="lattice"):
|
||||
if flavor == "lattice":
|
||||
if flavor in ["lattice", "lattice_ocr"]:
|
||||
for key in kwargs.keys():
|
||||
if key in stream_kwargs:
|
||||
kwargs.pop(key)
|
||||
@@ -365,7 +353,7 @@ def text_in_bbox(bbox, text):
|
||||
Returns
|
||||
-------
|
||||
t_bbox : list
|
||||
List of PDFMiner text objects that lie inside table.
|
||||
List of PDFMiner text objects that lie inside table, discarding the overlapping ones
|
||||
|
||||
"""
|
||||
lb = (bbox[0], bbox[1])
|
||||
@@ -376,7 +364,97 @@ def text_in_bbox(bbox, text):
|
||||
if lb[0] - 2 <= (t.x0 + t.x1) / 2.0 <= rt[0] + 2
|
||||
and lb[1] - 2 <= (t.y0 + t.y1) / 2.0 <= rt[1] + 2
|
||||
]
|
||||
return t_bbox
|
||||
|
||||
# Avoid duplicate text by discarding overlapping boxes
|
||||
rest = {t for t in t_bbox}
|
||||
for ba in t_bbox:
|
||||
for bb in rest.copy():
|
||||
if ba == bb:
|
||||
continue
|
||||
if bbox_intersect(ba, bb):
|
||||
# if the intersection is larger than 80% of ba's size, we keep the longest
|
||||
if (bbox_intersection_area(ba, bb) / bbox_area(ba)) > 0.8:
|
||||
if bbox_longer(bb, ba):
|
||||
rest.discard(ba)
|
||||
unique_boxes = list(rest)
|
||||
|
||||
return unique_boxes
|
||||
|
||||
|
||||
def bbox_intersection_area(ba, bb) -> float:
|
||||
"""Returns area of the intersection of the bounding boxes of two PDFMiner objects.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
ba : PDFMiner text object
|
||||
bb : PDFMiner text object
|
||||
|
||||
Returns
|
||||
-------
|
||||
intersection_area : float
|
||||
Area of the intersection of the bounding boxes of both objects
|
||||
|
||||
"""
|
||||
x_left = max(ba.x0, bb.x0)
|
||||
y_top = min(ba.y1, bb.y1)
|
||||
x_right = min(ba.x1, bb.x1)
|
||||
y_bottom = max(ba.y0, bb.y0)
|
||||
|
||||
if x_right < x_left or y_bottom > y_top:
|
||||
return 0.0
|
||||
|
||||
intersection_area = (x_right - x_left) * (y_top - y_bottom)
|
||||
return intersection_area
|
||||
|
||||
|
||||
def bbox_area(bb) -> float:
|
||||
"""Returns area of the bounding box of a PDFMiner object.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
bb : PDFMiner text object
|
||||
|
||||
Returns
|
||||
-------
|
||||
area : float
|
||||
Area of the bounding box of the object
|
||||
|
||||
"""
|
||||
return (bb.x1 - bb.x0) * (bb.y1 - bb.y0)
|
||||
|
||||
|
||||
def bbox_intersect(ba, bb) -> bool:
|
||||
"""Returns True if the bounding boxes of two PDFMiner objects intersect.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
ba : PDFMiner text object
|
||||
bb : PDFMiner text object
|
||||
|
||||
Returns
|
||||
-------
|
||||
overlaps : bool
|
||||
True if the bounding boxes intersect
|
||||
|
||||
"""
|
||||
return ba.x1 >= bb.x0 and bb.x1 >= ba.x0 and ba.y1 >= bb.y0 and bb.y1 >= ba.y0
|
||||
|
||||
|
||||
def bbox_longer(ba, bb) -> bool:
|
||||
"""Returns True if the bounding box of the first PDFMiner object is longer or equal to the second.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
ba : PDFMiner text object
|
||||
bb : PDFMiner text object
|
||||
|
||||
Returns
|
||||
-------
|
||||
longer : bool
|
||||
True if the bounding box of the first object is longer or equal
|
||||
|
||||
"""
|
||||
return (ba.x1 - ba.x0) >= (bb.x1 - bb.x0)
|
||||
|
||||
|
||||
def merge_close_lines(ar, line_tol=2):
|
||||
@@ -423,7 +501,7 @@ def text_strip(text, strip=""):
|
||||
return text
|
||||
|
||||
stripped = re.sub(
|
||||
r"[{}]".format("".join(map(re.escape, strip))), "", text, re.UNICODE
|
||||
fr"[{''.join(map(re.escape, strip))}]", "", text, flags=re.UNICODE
|
||||
)
|
||||
return stripped
|
||||
|
||||
@@ -660,9 +738,7 @@ def get_table_index(
|
||||
text_range = (t.x0, t.x1)
|
||||
col_range = (table.cols[0][0], table.cols[-1][1])
|
||||
warnings.warn(
|
||||
"{} {} does not lie in column range {}".format(
|
||||
text, text_range, col_range
|
||||
)
|
||||
f"{text} {text_range} does not lie in column range {col_range}"
|
||||
)
|
||||
r_idx = r
|
||||
c_idx = lt_col_overlap.index(max(lt_col_overlap))
|
||||
@@ -794,7 +870,7 @@ def get_page_layout(
|
||||
parser = PDFParser(f)
|
||||
document = PDFDocument(parser)
|
||||
if not document.is_extractable:
|
||||
raise PDFTextExtractionNotAllowed
|
||||
raise PDFTextExtractionNotAllowed(f"Text extraction is not allowed: {filename}")
|
||||
laparams = LAParams(
|
||||
char_margin=char_margin,
|
||||
line_margin=line_margin,
|
||||
|
||||
@@ -63,7 +63,7 @@ master_doc = 'index'
|
||||
|
||||
# General information about the project.
|
||||
project = u'Camelot'
|
||||
copyright = u'2019, Camelot Developers'
|
||||
copyright = u'2020, Camelot Developers'
|
||||
author = u'Vinayak Mehta'
|
||||
|
||||
# The version info for the project you're documenting, acts as replacement for
|
||||
|
||||
@@ -37,7 +37,7 @@ Setting up a development environment
|
||||
|
||||
To install the dependencies needed for development, you can use pip::
|
||||
|
||||
$ pip install camelot-py[dev]
|
||||
$ pip install "camelot-py[dev]"
|
||||
|
||||
Alternatively, you can clone the project repository, and install using pip::
|
||||
|
||||
|
||||
@@ -33,15 +33,18 @@ Release v\ |version|. (:ref:`Installation <install>`)
|
||||
.. image:: https://img.shields.io/badge/code%20style-black-000000.svg
|
||||
:target: https://github.com/ambv/black
|
||||
|
||||
**Camelot** is a Python library that makes it easy for *anyone* to extract tables from PDF files!
|
||||
.. image:: https://img.shields.io/badge/continous%20quality-deepsource-lightgrey
|
||||
:target: https://deepsource.io/gh/camelot-dev/camelot/?ref=repository-badge
|
||||
|
||||
.. note:: You can also check out `Excalibur`_, which is a web interface for Camelot!
|
||||
**Camelot** is a Python library that can help you extract tables from PDFs!
|
||||
|
||||
.. note:: You can also check out `Excalibur`_, the web interface to Camelot!
|
||||
|
||||
.. _Excalibur: https://github.com/camelot-dev/excalibur
|
||||
|
||||
----
|
||||
|
||||
**Here's how you can extract tables from PDF files.** Check out the PDF used in this example `here`_.
|
||||
**Here's how you can extract tables from PDFs.** You can check out the PDF used in this example `here`_.
|
||||
|
||||
.. _here: _static/pdf/foo.pdf
|
||||
|
||||
@@ -67,7 +70,7 @@ Release v\ |version|. (:ref:`Installation <install>`)
|
||||
.. csv-table::
|
||||
:file: _static/csv/foo.csv
|
||||
|
||||
There's a :ref:`command-line interface <cli>` too!
|
||||
Camelot also comes packaged with a :ref:`command-line interface <cli>`!
|
||||
|
||||
.. note:: Camelot only works with text-based PDFs and not scanned documents. (As Tabula `explains`_, "If you can click and drag to select text in your table in a PDF viewer, then your PDF is text-based".)
|
||||
|
||||
@@ -76,27 +79,27 @@ There's a :ref:`command-line interface <cli>` too!
|
||||
Why Camelot?
|
||||
------------
|
||||
|
||||
- **You are in control.** Unlike other libraries and tools which either give a nice output or fail miserably (with no in-between), Camelot gives you the power to tweak table extraction. (This is important since everything in the real world, including PDF table extraction, is fuzzy.)
|
||||
- *Bad* tables can be discarded based on **metrics** like accuracy and whitespace, without ever having to manually look at each table.
|
||||
- Each table is a **pandas DataFrame**, which seamlessly integrates into `ETL and data analysis workflows`_.
|
||||
- **Export** to multiple formats, including JSON, Excel and HTML.
|
||||
|
||||
See `comparison with other PDF table extraction libraries and tools`_.
|
||||
- **Configurability**: Camelot gives you control over the table extraction process with its :ref:`tweakable settings <advanced>`.
|
||||
- **Metrics**: Bad tables can be discarded based on metrics like accuracy and whitespace, without having to manually look at each table.
|
||||
- **Output**: Each table is extracted into a **pandas DataFrame**, which seamlessly integrates into `ETL and data analysis workflows`_. You can also export tables to multiple formats, which include CSV, JSON, Excel, HTML and Sqlite.
|
||||
|
||||
.. _ETL and data analysis workflows: https://gist.github.com/vinayak-mehta/e5949f7c2410a0e12f25d3682dc9e873
|
||||
.. _comparison with other PDF table extraction libraries and tools: https://github.com/camelot-dev/camelot/wiki/Comparison-with-other-PDF-Table-Extraction-libraries-and-tools
|
||||
|
||||
Support us on OpenCollective
|
||||
----------------------------
|
||||
See `comparison with similar libraries and tools`_.
|
||||
|
||||
If Camelot helped you extract tables from PDFs, please consider supporting its development by `becoming a backer or a sponsor on OpenCollective`_!
|
||||
.. _comparison with similar libraries and tools: https://github.com/camelot-dev/camelot/wiki/Comparison-with-other-PDF-Table-Extraction-libraries-and-tools
|
||||
|
||||
.. _becoming a backer or a sponsor on OpenCollective: https://opencollective.com/camelot
|
||||
Support the development
|
||||
-----------------------
|
||||
|
||||
If Camelot has helped you, please consider supporting its development with a one-time or monthly donation `on OpenCollective`_!
|
||||
|
||||
.. _on OpenCollective: https://opencollective.com/camelot
|
||||
|
||||
The User Guide
|
||||
--------------
|
||||
|
||||
This part of the documentation begins with some background information about why Camelot was created, takes a small dip into the implementation details and then focuses on step-by-step instructions for getting the most out of Camelot.
|
||||
This part of the documentation begins with some background information about why Camelot was created, takes you through some implementation details, and then focuses on step-by-step instructions for getting the most out of Camelot.
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
@@ -112,8 +115,7 @@ This part of the documentation begins with some background information about why
|
||||
The API Documentation/Guide
|
||||
---------------------------
|
||||
|
||||
If you are looking for information on a specific function, class, or method,
|
||||
this part of the documentation is for you.
|
||||
If you are looking for information on a specific function, class, or method, this part of the documentation is for you.
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
@@ -123,8 +125,7 @@ this part of the documentation is for you.
|
||||
The Contributor Guide
|
||||
---------------------
|
||||
|
||||
If you want to contribute to the project, this part of the documentation is for
|
||||
you.
|
||||
If you want to contribute to the project, this part of the documentation is for you.
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
|
||||
@@ -66,8 +66,7 @@ Let's plot all the text present on the table's PDF page.
|
||||
|
||||
::
|
||||
|
||||
>>> camelot.plot(tables[0], kind='text')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='text').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -93,8 +92,7 @@ Let's plot the table (to see if it was detected correctly or not). This plot typ
|
||||
|
||||
::
|
||||
|
||||
>>> camelot.plot(tables[0], kind='grid')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='grid').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -118,8 +116,7 @@ Now, let's plot all table boundaries present on the table's PDF page.
|
||||
|
||||
::
|
||||
|
||||
>>> camelot.plot(tables[0], kind='contour')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='contour').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -141,8 +138,7 @@ Cool, let's plot all line segments present on the table's PDF page.
|
||||
|
||||
::
|
||||
|
||||
>>> camelot.plot(tables[0], kind='line')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='line').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -164,8 +160,7 @@ Finally, let's plot all line intersections present on the table's PDF page.
|
||||
|
||||
::
|
||||
|
||||
>>> camelot.plot(tables[0], kind='joint')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='joint').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -187,8 +182,7 @@ You can also visualize the textedges found on a page by specifying ``kind='texte
|
||||
|
||||
::
|
||||
|
||||
>>> camelot.plot(tables[0], kind='textedge')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='textedge').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -375,8 +369,7 @@ Let's see the table area that is detected by default.
|
||||
::
|
||||
|
||||
>>> tables = camelot.read_pdf('edge_tol.pdf', flavor='stream')
|
||||
>>> camelot.plot(tables[0], kind='contour')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='contour').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -396,8 +389,7 @@ To improve the detected area, you can increase the ``edge_tol`` (default: 50) va
|
||||
::
|
||||
|
||||
>>> tables = camelot.read_pdf('edge_tol.pdf', flavor='stream', edge_tol=500)
|
||||
>>> camelot.plot(tables[0], kind='contour')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='contour').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
@@ -472,8 +464,7 @@ Let's plot the table for this PDF.
|
||||
::
|
||||
|
||||
>>> tables = camelot.read_pdf('short_lines.pdf')
|
||||
>>> camelot.plot(tables[0], kind='grid')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='grid').show()
|
||||
|
||||
.. figure:: ../_static/png/short_lines_1.png
|
||||
:alt: A plot of the PDF table with short lines
|
||||
@@ -484,8 +475,7 @@ Clearly, the smaller lines separating the headers, couldn't be detected. Let's t
|
||||
::
|
||||
|
||||
>>> tables = camelot.read_pdf('short_lines.pdf', line_scale=40)
|
||||
>>> camelot.plot(tables[0], kind='grid')
|
||||
>>> plt.show()
|
||||
>>> camelot.plot(tables[0], kind='grid').show()
|
||||
|
||||
.. tip::
|
||||
Here's how you can do the same with the :ref:`command-line interface <cli>`.
|
||||
|
||||
@@ -16,11 +16,11 @@ Stream can be used to parse tables that have whitespaces between cells to simula
|
||||
|
||||
1. Words on the PDF page are grouped into text rows based on their *y* axis overlaps.
|
||||
|
||||
2. Textedges are calculated and then used to guess interesting table areas on the PDF page. You can read `Anssi Nurminen's master's thesis <http://dspace.cc.tut.fi/dpub/bitstream/handle/123456789/21520/Nurminen.pdf?sequence=3>`_ to know more about this table detection technique. [See pages 20, 35 and 40]
|
||||
2. Textedges are calculated and then used to guess interesting table areas on the PDF page. You can read `Anssi Nurminen's master's thesis <https://pdfs.semanticscholar.org/a9b1/67a86fb189bfcd366c3839f33f0404db9c10.pdf>`_ to know more about this table detection technique. [See pages 20, 35 and 40]
|
||||
|
||||
3. The number of columns inside each table area are then guessed. This is done by calculating the mode of number of words in each text row. Based on this mode, words in each text row are chosen to calculate a list of column *x* ranges.
|
||||
|
||||
4. Words that lie inside/outside the current column *x* ranges are then used to extend extend the current list of columns.
|
||||
4. Words that lie inside/outside the current column *x* ranges are then used to extend the current list of columns.
|
||||
|
||||
5. Finally, a table is formed using the text rows' *y* ranges and column *x* ranges and words found on the page are assigned to the table's cells based on their *x* and *y* coordinates.
|
||||
|
||||
|
||||
@@ -3,72 +3,59 @@
|
||||
Installation of dependencies
|
||||
============================
|
||||
|
||||
The dependencies `Tkinter`_ and `ghostscript`_ can be installed using your system's package manager. You can run one of the following, based on your OS.
|
||||
|
||||
.. _Tkinter: https://wiki.python.org/moin/TkInter
|
||||
.. _ghostscript: https://www.ghostscript.com
|
||||
The dependencies `Ghostscript <https://www.ghostscript.com>`_ and `Tkinter <https://wiki.python.org/moin/TkInter>`_ can be installed using your system's package manager or by running their installer.
|
||||
|
||||
OS-specific instructions
|
||||
------------------------
|
||||
|
||||
For Ubuntu
|
||||
^^^^^^^^^^
|
||||
Ubuntu
|
||||
^^^^^^
|
||||
::
|
||||
|
||||
$ apt install python-tk ghostscript
|
||||
$ apt install ghostscript python3-tk
|
||||
|
||||
Or for Python 3::
|
||||
|
||||
$ apt install python3-tk ghostscript
|
||||
|
||||
For macOS
|
||||
^^^^^^^^^
|
||||
MacOS
|
||||
^^^^^
|
||||
::
|
||||
|
||||
$ brew install tcl-tk ghostscript
|
||||
$ brew install ghostscript tcl-tk
|
||||
|
||||
For Windows
|
||||
^^^^^^^^^^^
|
||||
Windows
|
||||
^^^^^^^
|
||||
|
||||
For Tkinter, you can download the `ActiveTcl Community Edition`_ from ActiveState. For ghostscript, you can get the installer at the `ghostscript downloads page`_.
|
||||
For Ghostscript, you can get the installer at their `downloads page <https://www.ghostscript.com/download/gsdnld.html>`_. And for Tkinter, you can download the `ActiveTcl Community Edition <https://www.activestate.com/activetcl/downloads>`_ from ActiveState.
|
||||
|
||||
.. _ActiveTcl Community Edition: https://www.activestate.com/activetcl/downloads
|
||||
.. _ghostscript downloads page: https://www.ghostscript.com/download/gsdnld.html
|
||||
.. _as shown here: https://java.com/en/download/help/path.xml
|
||||
Checks to see if dependencies are installed correctly
|
||||
-----------------------------------------------------
|
||||
|
||||
Checks to see if dependencies were installed correctly
|
||||
------------------------------------------------------
|
||||
You can run the following checks to see if the dependencies were installed correctly.
|
||||
|
||||
You can do the following checks to see if the dependencies were installed correctly.
|
||||
For Ghostscript
|
||||
^^^^^^^^^^^^^^^
|
||||
|
||||
Open the Python REPL and run the following:
|
||||
|
||||
For Ubuntu/MacOS::
|
||||
|
||||
>>> from ctypes.util import find_library
|
||||
>>> find_library("gs")
|
||||
"libgs.so.9"
|
||||
|
||||
For Windows::
|
||||
|
||||
>>> from ctypes.util import find_library
|
||||
>>> find_library("".join(("gsdll", str(ctypes.sizeof(ctypes.c_voidp) * 8), ".dll"))
|
||||
<name-of-ghostscript-library-on-windows>
|
||||
|
||||
**Check:** The output of the ``find_library`` function should not be empty.
|
||||
|
||||
If the output is empty, then it's possible that the Ghostscript library is not available one of the ``LD_LIBRARY_PATH``/``DYLD_LIBRARY_PATH``/``PATH`` variables depending on your operating system. In this case, you may have to modify one of those path variables.
|
||||
|
||||
For Tkinter
|
||||
^^^^^^^^^^^
|
||||
|
||||
Launch Python, and then at the prompt, type::
|
||||
|
||||
>>> import Tkinter
|
||||
|
||||
Or in Python 3::
|
||||
Launch Python and then import Tkinter::
|
||||
|
||||
>>> import tkinter
|
||||
|
||||
If you have Tkinter, Python will not print an error message, and if not, you will see an ``ImportError``.
|
||||
|
||||
For ghostscript
|
||||
^^^^^^^^^^^^^^^
|
||||
|
||||
Run the following to check the ghostscript version.
|
||||
|
||||
For Ubuntu/macOS::
|
||||
|
||||
$ gs -version
|
||||
|
||||
For Windows::
|
||||
|
||||
C:\> gswin64c.exe -version
|
||||
|
||||
Or for Windows 32-bit::
|
||||
|
||||
C:\> gswin32c.exe -version
|
||||
|
||||
If you have ghostscript, you should see the ghostscript version and copyright information.
|
||||
**Check:** Importing ``tkinter`` should not raise an import error.
|
||||
|
||||
@@ -5,42 +5,35 @@ Installation of Camelot
|
||||
|
||||
This part of the documentation covers the steps to install Camelot.
|
||||
|
||||
Using conda
|
||||
-----------
|
||||
After :ref:`installing the dependencies <install_deps>`, which include `Ghostscript <https://www.ghostscript.com>`_ and `Tkinter <https://wiki.python.org/moin/TkInter>`_, you can use one of the following methods to install Camelot:
|
||||
|
||||
The easiest way to install Camelot is to install it with `conda`_, which is a package manager and environment management system for the `Anaconda`_ distribution.
|
||||
::
|
||||
.. warning:: The ``lattice`` flavor will fail to run if Ghostscript is not installed. You may run into errors as shown in `issue #193 <https://github.com/camelot-dev/camelot/issues/193>`_.
|
||||
|
||||
pip
|
||||
---
|
||||
|
||||
To install Camelot from PyPI using ``pip``, please include the extra ``cv`` requirement as shown::
|
||||
|
||||
$ pip install "camelot-py[cv]"
|
||||
|
||||
conda
|
||||
-----
|
||||
|
||||
`conda`_ is a package manager and environment management system for the `Anaconda <https://anaconda.org>`_ distribution. It can be used to install Camelot from the ``conda-forge`` channel::
|
||||
|
||||
$ conda install -c conda-forge camelot-py
|
||||
|
||||
.. note:: Camelot is available for Python 2.7, 3.5, 3.6 and 3.7 on Linux, macOS and Windows. For Windows, you will need to install ghostscript which you can get from their `downloads page`_.
|
||||
|
||||
.. _conda: https://conda.io/docs/
|
||||
.. _Anaconda: http://docs.continuum.io/anaconda/
|
||||
.. _downloads page: https://www.ghostscript.com/download/gsdnld.html
|
||||
.. _conda-forge: https://conda-forge.org/
|
||||
|
||||
Using pip
|
||||
---------
|
||||
|
||||
After :ref:`installing the dependencies <install_deps>`, which include `Tkinter`_ and `ghostscript`_, you can simply use pip to install Camelot::
|
||||
|
||||
$ pip install camelot-py[cv]
|
||||
|
||||
.. _Tkinter: https://wiki.python.org/moin/TkInter
|
||||
.. _ghostscript: https://www.ghostscript.com
|
||||
|
||||
From the source code
|
||||
--------------------
|
||||
|
||||
After :ref:`installing the dependencies <install_deps>`, you can install from the source by:
|
||||
After :ref:`installing the dependencies <install_deps>`, you can install Camelot from source by:
|
||||
|
||||
1. Cloning the GitHub repository.
|
||||
::
|
||||
|
||||
$ git clone https://www.github.com/camelot-dev/camelot
|
||||
|
||||
2. Then simply using pip again.
|
||||
2. And then simply using pip again.
|
||||
::
|
||||
|
||||
$ cd camelot
|
||||
|
||||
@@ -1,8 +0,0 @@
|
||||
click>=6.7
|
||||
matplotlib>=2.2.3
|
||||
numpy>=1.13.3
|
||||
opencv-python>=3.4.2.17
|
||||
openpyxl>=2.5.8
|
||||
pandas>=0.23.4
|
||||
pdfminer.six>=20170720
|
||||
PyPDF2>=1.26.0
|
||||
@@ -19,7 +19,7 @@ requires = [
|
||||
'numpy>=1.13.3',
|
||||
'openpyxl>=2.5.8',
|
||||
'pandas>=0.23.4',
|
||||
'pdfminer.six>=20170720',
|
||||
'pdfminer.six>=20200726',
|
||||
'PyPDF2>=1.26.0'
|
||||
]
|
||||
|
||||
@@ -27,20 +27,24 @@ cv_requires = [
|
||||
'opencv-python>=3.4.2.17'
|
||||
]
|
||||
|
||||
ocr_requires = [
|
||||
'easyocr>=1.1.10'
|
||||
]
|
||||
|
||||
plot_requires = [
|
||||
'matplotlib>=2.2.3',
|
||||
]
|
||||
|
||||
dev_requires = [
|
||||
'codecov>=2.0.15',
|
||||
'pytest>=3.8.0',
|
||||
'pytest-cov>=2.6.0',
|
||||
'pytest-mpl>=0.10',
|
||||
'pytest-runner>=4.2',
|
||||
'Sphinx>=1.7.9'
|
||||
'pytest>=5.4.3',
|
||||
'pytest-cov>=2.10.0',
|
||||
'pytest-mpl>=0.11',
|
||||
'pytest-runner>=5.2',
|
||||
'Sphinx>=3.1.2'
|
||||
]
|
||||
|
||||
all_requires = cv_requires + plot_requires
|
||||
all_requires = cv_requires + ocr_requires + plot_requires
|
||||
dev_requires = dev_requires + all_requires
|
||||
|
||||
|
||||
@@ -57,10 +61,11 @@ def setup_package():
|
||||
packages=find_packages(exclude=('tests',)),
|
||||
install_requires=requires,
|
||||
extras_require={
|
||||
'all': all_requires,
|
||||
'cv': cv_requires,
|
||||
'ocr': ocr_requires,
|
||||
'plot': plot_requires,
|
||||
'all': all_requires,
|
||||
'dev': dev_requires,
|
||||
'plot': plot_requires
|
||||
},
|
||||
entry_points={
|
||||
'console_scripts': [
|
||||
@@ -71,10 +76,9 @@ def setup_package():
|
||||
# Trove classifiers
|
||||
# Full list: https://pypi.python.org/pypi?%3Aaction=list_classifiers
|
||||
'License :: OSI Approved :: MIT License',
|
||||
'Programming Language :: Python :: 2.7',
|
||||
'Programming Language :: Python :: 3.5',
|
||||
'Programming Language :: Python :: 3.6',
|
||||
'Programming Language :: Python :: 3.7'
|
||||
'Programming Language :: Python :: 3.7',
|
||||
'Programming Language :: Python :: 3.8'
|
||||
])
|
||||
|
||||
try:
|
||||
|
||||
@@ -1,2 +1,3 @@
|
||||
import matplotlib
|
||||
matplotlib.use('agg')
|
||||
|
||||
matplotlib.use("agg")
|
||||
|
||||
@@ -1,19 +1,7 @@
|
||||
# -*- coding: utf-8 -*-
|
||||
|
||||
from __future__ import unicode_literals
|
||||
|
||||
|
||||
data_stream = [
|
||||
[
|
||||
"",
|
||||
"Table: 5 Public Health Outlay 2012-13 (Budget Estimates) (Rs. in 000)",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
],
|
||||
["States-A", "Revenue", "", "Capital", "", "Total", "Others(1)", "Total"],
|
||||
["", "", "", "", "", "Revenue &", "", ""],
|
||||
["", "Medical &", "Family", "Medical &", "Family", "", "", ""],
|
||||
@@ -829,18 +817,6 @@ data_stream_table_rotated = [
|
||||
]
|
||||
|
||||
data_stream_two_tables_1 = [
|
||||
[
|
||||
"[In thousands (11,062.6 represents 11,062,600) For year ending December 31. Based on Uniform Crime Reporting (UCR)",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
],
|
||||
[
|
||||
"Program. Represents arrests reported (not charged) by 12,910 agencies with a total population of 247,526,916 as estimated",
|
||||
"",
|
||||
@@ -915,7 +891,7 @@ data_stream_two_tables_1 = [
|
||||
"2,330 .9",
|
||||
],
|
||||
[
|
||||
"Violent crime . . . . . . . .\n . .\n . .\n . .\n . .\n . .",
|
||||
"Violent crime . . . . . . . .\n . .\n . .\n . .\n . .\n . .",
|
||||
"467 .9",
|
||||
"69 .1",
|
||||
"398 .8",
|
||||
@@ -1300,29 +1276,10 @@ data_stream_two_tables_1 = [
|
||||
"",
|
||||
"",
|
||||
],
|
||||
[
|
||||
"",
|
||||
"Source: U.S. Department of Justice, Federal Bureau of Investigation, Uniform Crime Reports, Arrests Master Files.",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
],
|
||||
]
|
||||
|
||||
|
||||
data_stream_two_tables_2 = [
|
||||
[
|
||||
"",
|
||||
"Source: U.S. Department of Justice, Federal Bureau of Investigation, Uniform Crime Reports, Arrests Master Files.",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
],
|
||||
["Table 325. Arrests by Race: 2009", "", "", "", "", ""],
|
||||
[
|
||||
"[Based on Uniform Crime Reporting (UCR) Program. Represents arrests reported (not charged) by 12,371 agencies",
|
||||
@@ -1352,7 +1309,7 @@ data_stream_two_tables_2 = [
|
||||
"123,656",
|
||||
],
|
||||
[
|
||||
"Violent crime . . . . . . . .\n . .\n . .\n . .\n . .\n .\n .\n . .\n . .\n .\n .\n .\n .\n . .",
|
||||
"Violent crime . . . . . . . .\n . .\n . .\n . .\n . .\n .\n .\n . .\n . .\n .\n .\n .\n .\n . .",
|
||||
"456,965",
|
||||
"268,346",
|
||||
"177,766",
|
||||
@@ -1600,16 +1557,9 @@ data_stream_two_tables_2 = [
|
||||
"3,950",
|
||||
],
|
||||
["1 Except forcible rape and prostitution.", "", "", "", "", ""],
|
||||
[
|
||||
"",
|
||||
"Source: U.S. Department of Justice, Federal Bureau of Investigation, “Crime in the United States, Arrests,” September 2010,",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
],
|
||||
]
|
||||
|
||||
|
||||
data_stream_table_areas = [
|
||||
["", "One Withholding"],
|
||||
["Payroll Period", "Allowance"],
|
||||
@@ -1776,18 +1726,7 @@ data_stream_columns = [
|
||||
]
|
||||
|
||||
data_stream_split_text = [
|
||||
[
|
||||
"FEB",
|
||||
"RUAR",
|
||||
"Y 2014 M27 (BUS)",
|
||||
"",
|
||||
"ALPHABETIC LISTING BY T",
|
||||
"YPE",
|
||||
"",
|
||||
"",
|
||||
"",
|
||||
"ABLPDM27",
|
||||
],
|
||||
["FEB", "RUAR", "Y 2014 M27 (BUS)", "", "", "", "", "", "", ""],
|
||||
["", "", "", "", "OF ACTIVE LICENSES", "", "", "", "", "3/19/2014"],
|
||||
["", "", "", "", "OKLAHOMA ABLE COMMIS", "SION", "", "", "", ""],
|
||||
["LICENSE", "", "", "", "PREMISE", "", "", "", "", ""],
|
||||
@@ -2121,6 +2060,7 @@ data_stream_split_text = [
|
||||
],
|
||||
]
|
||||
|
||||
|
||||
data_stream_flag_size = [
|
||||
[
|
||||
"States",
|
||||
@@ -2820,7 +2760,7 @@ data_arabic = [
|
||||
]
|
||||
|
||||
data_stream_layout_kwargs = [
|
||||
["V i n s a u Ve r r e", ""],
|
||||
["V i n s a u V e r r e", ""],
|
||||
["Les Blancs", "12.5CL"],
|
||||
["A.O.P Côtes du Rhône", ""],
|
||||
["Domaine de la Guicharde « Autour de la chapelle » 2016", "8 €"],
|
||||
@@ -2858,3 +2798,51 @@ data_stream_layout_kwargs = [
|
||||
["A.O.P Cornas", ""],
|
||||
["Domaine Lionnet « Terre Brûlée » 2012", "15 €"],
|
||||
]
|
||||
|
||||
data_stream_duplicated_text = [
|
||||
['', '2012 BETTER VARIETIES Harvest Report for Minnesota Central [ MNCE ]', '', '', '', '', '', '', '', '',
|
||||
'ALL SEASON TEST'],
|
||||
['', 'Doug Toreen, Renville County, MN 55310 [ BIRD ISLAND ]', '', '', '', '', '', '', '', '',
|
||||
'1.3 - 2.0 MAT. GROUP'],
|
||||
['PREV. CROP/HERB:', 'Corn / Surpass, Roundup', '', '', '', '', '', '', '', '', 'S2MNCE01'],
|
||||
['SOIL DESCRIPTION:', '', 'Canisteo clay loam, mod. well drained, non-irrigated', '', '', '', '', '', '', '', ''],
|
||||
['SOIL CONDITIONS:', '', 'High P, high K, 6.7 pH, 3.9% OM, Low SCN', '', '', '', '', '', '', '', '30" ROW SPACING'],
|
||||
['TILLAGE/CULTIVATION:', 'conventional w/ fall till', '', '', '', '', '', '', '', '', ''],
|
||||
['PEST MANAGEMENT:', 'Roundup twice', '', '', '', '', '', '', '', '', ''],
|
||||
['SEEDED - RATE:', 'May 15', '140 000 /A', '', '', '', '', '', '', 'TOP 30 for YIELD of 63 TESTED', ''],
|
||||
['HARVESTED - STAND:', 'Oct 3', '122 921 /A', '', '', '', '', '', '', 'AVERAGE of (3) REPLICATIONS', ''],
|
||||
['', '', '', '', 'SCN', 'Seed', 'Yield', 'Moisture', 'Lodging', 'Stand', 'Gross'],
|
||||
['Company/Brand', 'Product/Brand†', 'Technol.†', 'Mat.', 'Resist.', 'Trmt.†', 'Bu/A', '%', '%', '(x 1000)',
|
||||
'Income'], ['Kruger', 'K2 1901', 'RR2Y', '1.9', 'R', 'Ac,PV', '56.4', '7.6', '0', '126.3', '$846'],
|
||||
['Stine', '19RA02 §', 'RR2Y', '1.9', 'R', 'CMB', '55.3', '7.6', '0', '120.0', '$830'],
|
||||
['Wensman', 'W 3190NR2', 'RR2Y', '1.9', 'R', 'Ac', '54.5', '7.6', '0', '119.5', '$818'],
|
||||
['Hefty', 'H17Y12', 'RR2Y', '1.7', 'MR', 'I', '53.7', '7.7', '0', '124.4', '$806'],
|
||||
['Dyna-Gro', 'S15RY53', 'RR2Y', '1.5', 'R', 'Ac', '53.6', '7.7', '0', '126.8', '$804'],
|
||||
['LG Seeds', 'C2050R2', 'RR2Y', '2.1', 'R', 'Ac', '53.6', '7.7', '0', '123.9', '$804'],
|
||||
['Titan Pro', '19M42', 'RR2Y', '1.9', 'R', 'CMB', '53.6', '7.7', '0', '121.0', '$804'],
|
||||
['Stine', '19RA02 (2) §', 'RR2Y', '1.9', 'R', 'CMB', '53.4', '7.7', '0', '123.9', '$801'],
|
||||
['Asgrow', 'AG1832 §', 'RR2Y', '1.8', 'MR', 'Ac,PV', '52.9', '7.7', '0', '122.0', '$794'],
|
||||
['Prairie Brand', 'PB-1566R2', 'RR2Y', '1.5', 'R', 'CMB', '52.8', '7.7', '0', '122.9', '$792'],
|
||||
['Channel', '1901R2', 'RR2Y', '1.9', 'R', 'Ac,PV', '52.8', '7.6', '0', '123.4', '$791'],
|
||||
['Titan Pro', '20M1', 'RR2Y', '2.0', 'R', 'Am', '52.5', '7.5', '0', '124.4', '$788'],
|
||||
['Kruger', 'K2-2002', 'RR2Y', '2.0', 'R', 'Ac,PV', '52.4', '7.9', '0', '125.4', '$786'],
|
||||
['Channel', '1700R2', 'RR2Y', '1.7', 'R', 'Ac,PV', '52.3', '7.9', '0', '123.9', '$784'],
|
||||
['Hefty', 'H16Y11', 'RR2Y', '1.6', 'MR', 'I', '51.4', '7.6', '0', '123.9', '$771'],
|
||||
['Anderson', '162R2Y', 'RR2Y', '1.6', 'R', 'None', '51.3', '7.5', '0', '119.5', '$770'],
|
||||
['Titan Pro', '15M22', 'RR2Y', '1.5', 'R', 'CMB', '51.3', '7.8', '0', '125.4', '$769'],
|
||||
['Dairyland', 'DSR-1710R2Y', 'RR2Y', '1.7', 'R', 'CMB', '51.3', '7.7', '0', '122.0', '$769'],
|
||||
['Hefty', 'H20R3', 'RR2Y', '2.0', 'MR', 'I', '50.5', '8.2', '0', '121.0', '$757'],
|
||||
['Prairie Brand', 'PB 1743R2', 'RR2Y', '1.7', 'R', 'CMB', '50.2', '7.7', '0', '125.8', '$752'],
|
||||
['Gold Country', '1741', 'RR2Y', '1.7', 'R', 'Ac', '50.1', '7.8', '0', '123.9', '$751'],
|
||||
['Trelay', '20RR43', 'RR2Y', '2.0', 'R', 'Ac,Ex', '49.9', '7.6', '0', '127.8', '$749'],
|
||||
['Hefty', 'H14R3', 'RR2Y', '1.4', 'MR', 'I', '49.7', '7.7', '0', '122.9', '$746'],
|
||||
['Prairie Brand', 'PB-2099NRR2', 'RR2Y', '2.0', 'R', 'CMB', '49.6', '7.8', '0', '126.3', '$743'],
|
||||
['Wensman', 'W 3174NR2', 'RR2Y', '1.7', 'R', 'Ac', '49.3', '7.6', '0', '122.5', '$740'],
|
||||
['Kruger', 'K2 1602', 'RR2Y', '1.6', 'R', 'Ac,PV', '48.7', '7.6', '0', '125.4', '$731'],
|
||||
['NK Brand', 'S18-C2 §', 'RR2Y', '1.8', 'R', 'CMB', '48.7', '7.7', '0', '126.8', '$731'],
|
||||
['Kruger', 'K2 1902', 'RR2Y', '1.9', 'R', 'Ac,PV', '48.7', '7.5', '0', '124.4', '$730'],
|
||||
['Prairie Brand', 'PB-1823R2', 'RR2Y', '1.8', 'R', 'None', '48.5', '7.6', '0', '121.0', '$727'],
|
||||
['Gold Country', '1541', 'RR2Y', '1.5', 'R', 'Ac', '48.4', '7.6', '0', '110.4', '$726'],
|
||||
['', '', '', '', '', 'Test Average =', '47.6', '7.7', '0', '122.9', '$713'],
|
||||
['', '', '', '', '', 'LSD (0.10) =', '5.7', '0.3', 'ns', '37.8', '566.4']
|
||||
]
|
||||
|
||||
|
Before Width: | Height: | Size: 48 KiB After Width: | Height: | Size: 48 KiB |
|
Before Width: | Height: | Size: 6.7 KiB After Width: | Height: | Size: 6.7 KiB |
|
Before Width: | Height: | Size: 13 KiB After Width: | Height: | Size: 14 KiB |
|
Before Width: | Height: | Size: 8.8 KiB After Width: | Height: | Size: 8.9 KiB |
|
Before Width: | Height: | Size: 18 KiB After Width: | Height: | Size: 19 KiB |
@@ -114,31 +114,35 @@ def test_cli_password():
|
||||
def test_cli_output_format():
|
||||
with TemporaryDirectory() as tempdir:
|
||||
infile = os.path.join(testdir, "health.pdf")
|
||||
outfile = os.path.join(tempdir, "health.{}")
|
||||
|
||||
runner = CliRunner()
|
||||
|
||||
# json
|
||||
outfile = os.path.join(tempdir, "health.json")
|
||||
result = runner.invoke(
|
||||
cli,
|
||||
["--format", "json", "--output", outfile.format("json"), "stream", infile],
|
||||
["--format", "json", "--output", outfile, "stream", infile],
|
||||
)
|
||||
assert result.exit_code == 0
|
||||
|
||||
# excel
|
||||
outfile = os.path.join(tempdir, "health.xlsx")
|
||||
result = runner.invoke(
|
||||
cli,
|
||||
["--format", "excel", "--output", outfile.format("xlsx"), "stream", infile],
|
||||
["--format", "excel", "--output", outfile, "stream", infile],
|
||||
)
|
||||
assert result.exit_code == 0
|
||||
|
||||
# html
|
||||
outfile = os.path.join(tempdir, "health.html")
|
||||
result = runner.invoke(
|
||||
cli,
|
||||
["--format", "html", "--output", outfile.format("html"), "stream", infile],
|
||||
["--format", "html", "--output", outfile, "stream", infile],
|
||||
)
|
||||
assert result.exit_code == 0
|
||||
|
||||
# zip
|
||||
outfile = os.path.join(tempdir, "health.csv")
|
||||
result = runner.invoke(
|
||||
cli,
|
||||
[
|
||||
@@ -146,7 +150,7 @@ def test_cli_output_format():
|
||||
"--format",
|
||||
"csv",
|
||||
"--output",
|
||||
outfile.format("csv"),
|
||||
outfile,
|
||||
"stream",
|
||||
infile,
|
||||
],
|
||||
@@ -156,8 +160,8 @@ def test_cli_output_format():
|
||||
|
||||
def test_cli_quiet():
|
||||
with TemporaryDirectory() as tempdir:
|
||||
infile = os.path.join(testdir, "blank.pdf")
|
||||
outfile = os.path.join(tempdir, "blank.csv")
|
||||
infile = os.path.join(testdir, "empty.pdf")
|
||||
outfile = os.path.join(tempdir, "empty.csv")
|
||||
runner = CliRunner()
|
||||
|
||||
result = runner.invoke(
|
||||
|
||||
@@ -3,9 +3,11 @@
|
||||
import os
|
||||
|
||||
import pandas as pd
|
||||
from pandas.testing import assert_frame_equal
|
||||
|
||||
import camelot
|
||||
from camelot.core import Table, TableList
|
||||
from camelot.__version__ import generate_version
|
||||
|
||||
from .data import *
|
||||
|
||||
@@ -26,10 +28,10 @@ def test_password():
|
||||
|
||||
filename = os.path.join(testdir, "health_protected.pdf")
|
||||
tables = camelot.read_pdf(filename, password="ownerpass", flavor="stream")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
tables = camelot.read_pdf(filename, password="userpass", flavor="stream")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream():
|
||||
@@ -37,7 +39,7 @@ def test_stream():
|
||||
|
||||
filename = os.path.join(testdir, "health.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_table_rotated():
|
||||
@@ -45,11 +47,11 @@ def test_stream_table_rotated():
|
||||
|
||||
filename = os.path.join(testdir, "clockwise_table_2.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
filename = os.path.join(testdir, "anticlockwise_table_2.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_two_tables():
|
||||
@@ -71,7 +73,7 @@ def test_stream_table_regions():
|
||||
tables = camelot.read_pdf(
|
||||
filename, flavor="stream", table_regions=["320,460,573,335"]
|
||||
)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_table_areas():
|
||||
@@ -81,7 +83,7 @@ def test_stream_table_areas():
|
||||
tables = camelot.read_pdf(
|
||||
filename, flavor="stream", table_areas=["320,500,573,335"]
|
||||
)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_columns():
|
||||
@@ -91,7 +93,7 @@ def test_stream_columns():
|
||||
tables = camelot.read_pdf(
|
||||
filename, flavor="stream", columns=["67,180,230,425,475"], row_tol=10
|
||||
)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_split_text():
|
||||
@@ -104,7 +106,7 @@ def test_stream_split_text():
|
||||
columns=["72,95,209,327,442,529,566,606,683"],
|
||||
split_text=True,
|
||||
)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_flag_size():
|
||||
@@ -112,7 +114,7 @@ def test_stream_flag_size():
|
||||
|
||||
filename = os.path.join(testdir, "superscript.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream", flag_size=True)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_strip_text():
|
||||
@@ -120,7 +122,7 @@ def test_stream_strip_text():
|
||||
|
||||
filename = os.path.join(testdir, "detect_vertical_false.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream", strip_text=" ,\n")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_edge_tol():
|
||||
@@ -128,7 +130,7 @@ def test_stream_edge_tol():
|
||||
|
||||
filename = os.path.join(testdir, "edge_tol.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream", edge_tol=500)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_stream_layout_kwargs():
|
||||
@@ -138,7 +140,7 @@ def test_stream_layout_kwargs():
|
||||
tables = camelot.read_pdf(
|
||||
filename, flavor="stream", layout_kwargs={"detect_vertical": False}
|
||||
)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_lattice():
|
||||
@@ -148,7 +150,7 @@ def test_lattice():
|
||||
testdir, "tabula/icdar2013-dataset/competition-dataset-us/us-030.pdf"
|
||||
)
|
||||
tables = camelot.read_pdf(filename, pages="2")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_lattice_table_rotated():
|
||||
@@ -156,11 +158,11 @@ def test_lattice_table_rotated():
|
||||
|
||||
filename = os.path.join(testdir, "clockwise_table_1.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
filename = os.path.join(testdir, "anticlockwise_table_1.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_lattice_two_tables():
|
||||
@@ -179,7 +181,7 @@ def test_lattice_table_regions():
|
||||
|
||||
filename = os.path.join(testdir, "table_region.pdf")
|
||||
tables = camelot.read_pdf(filename, table_regions=["170,370,560,270"])
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_lattice_table_areas():
|
||||
@@ -187,7 +189,7 @@ def test_lattice_table_areas():
|
||||
|
||||
filename = os.path.join(testdir, "twotables_2.pdf")
|
||||
tables = camelot.read_pdf(filename, table_areas=["80,693,535,448"])
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_lattice_process_background():
|
||||
@@ -195,7 +197,7 @@ def test_lattice_process_background():
|
||||
|
||||
filename = os.path.join(testdir, "background_lines_1.pdf")
|
||||
tables = camelot.read_pdf(filename, process_background=True)
|
||||
assert df.equals(tables[1].df)
|
||||
assert_frame_equal(df, tables[1].df)
|
||||
|
||||
|
||||
def test_lattice_copy_text():
|
||||
@@ -203,7 +205,7 @@ def test_lattice_copy_text():
|
||||
|
||||
filename = os.path.join(testdir, "row_span_1.pdf")
|
||||
tables = camelot.read_pdf(filename, line_scale=60, copy_text="v")
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_lattice_shift_text():
|
||||
@@ -271,7 +273,7 @@ def test_arabic():
|
||||
|
||||
filename = os.path.join(testdir, "tabula/arabic.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
assert df.equals(tables[0].df)
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
|
||||
def test_table_order():
|
||||
@@ -297,3 +299,26 @@ def test_table_order():
|
||||
(1, 2),
|
||||
(1, 1),
|
||||
]
|
||||
|
||||
|
||||
def test_version_generation():
|
||||
version = (0, 7, 3)
|
||||
assert generate_version(version, prerelease=None, revision=None) == "0.7.3"
|
||||
|
||||
|
||||
def test_version_generation_with_prerelease_revision():
|
||||
version = (0, 7, 3)
|
||||
prerelease = "alpha"
|
||||
revision = 2
|
||||
assert (
|
||||
generate_version(version, prerelease=prerelease, revision=revision)
|
||||
== "0.7.3-alpha.2"
|
||||
)
|
||||
|
||||
|
||||
def test_stream_duplicated_text():
|
||||
df = pd.DataFrame(data_stream_duplicated_text)
|
||||
|
||||
filename = os.path.join(testdir, "birdisland.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
assert_frame_equal(df, tables[0].df)
|
||||
|
||||
@@ -10,88 +10,111 @@ import camelot
|
||||
|
||||
testdir = os.path.dirname(os.path.abspath(__file__))
|
||||
testdir = os.path.join(testdir, "files")
|
||||
filename = os.path.join(testdir, 'foo.pdf')
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
|
||||
|
||||
def test_unknown_flavor():
|
||||
message = ("Unknown flavor specified."
|
||||
" Use either 'lattice' or 'stream'")
|
||||
message = "Unknown flavor specified." " Use either 'lattice' or 'stream'"
|
||||
with pytest.raises(NotImplementedError, match=message):
|
||||
tables = camelot.read_pdf(filename, flavor='chocolate')
|
||||
tables = camelot.read_pdf(filename, flavor="chocolate")
|
||||
|
||||
|
||||
def test_input_kwargs():
|
||||
message = "columns cannot be used with flavor='lattice'"
|
||||
with pytest.raises(ValueError, match=message):
|
||||
tables = camelot.read_pdf(filename, columns=['10,20,30,40'])
|
||||
tables = camelot.read_pdf(filename, columns=["10,20,30,40"])
|
||||
|
||||
|
||||
def test_unsupported_format():
|
||||
message = 'File format not supported'
|
||||
filename = os.path.join(testdir, 'foo.csv')
|
||||
message = "File format not supported"
|
||||
filename = os.path.join(testdir, "foo.csv")
|
||||
with pytest.raises(NotImplementedError, match=message):
|
||||
tables = camelot.read_pdf(filename)
|
||||
|
||||
|
||||
def test_stream_equal_length():
|
||||
message = ("Length of table_areas and columns"
|
||||
" should be equal")
|
||||
message = "Length of table_areas and columns" " should be equal"
|
||||
with pytest.raises(ValueError, match=message):
|
||||
tables = camelot.read_pdf(filename, flavor='stream',
|
||||
table_areas=['10,20,30,40'], columns=['10,20,30,40', '10,20,30,40'])
|
||||
tables = camelot.read_pdf(
|
||||
filename,
|
||||
flavor="stream",
|
||||
table_areas=["10,20,30,40"],
|
||||
columns=["10,20,30,40", "10,20,30,40"],
|
||||
)
|
||||
|
||||
|
||||
def test_image_warning():
|
||||
filename = os.path.join(testdir, 'image.pdf')
|
||||
filename = os.path.join(testdir, "image.pdf")
|
||||
with warnings.catch_warnings():
|
||||
warnings.simplefilter('error')
|
||||
warnings.simplefilter("error")
|
||||
with pytest.raises(UserWarning) as e:
|
||||
tables = camelot.read_pdf(filename)
|
||||
assert str(e.value) == 'page-1 is image-based, camelot only works on text-based pages.'
|
||||
assert (
|
||||
str(e.value)
|
||||
== "page-1 is image-based, camelot only works on text-based pages."
|
||||
)
|
||||
|
||||
|
||||
def test_no_tables_found():
|
||||
filename = os.path.join(testdir, 'blank.pdf')
|
||||
def test_lattice_no_tables_on_page():
|
||||
filename = os.path.join(testdir, "empty.pdf")
|
||||
with warnings.catch_warnings():
|
||||
warnings.simplefilter('error')
|
||||
warnings.simplefilter("error")
|
||||
with pytest.raises(UserWarning) as e:
|
||||
tables = camelot.read_pdf(filename)
|
||||
assert str(e.value) == 'No tables found on page-1'
|
||||
tables = camelot.read_pdf(filename, flavor="lattice")
|
||||
assert str(e.value) == "No tables found on page-1"
|
||||
|
||||
|
||||
def test_stream_no_tables_on_page():
|
||||
filename = os.path.join(testdir, "empty.pdf")
|
||||
with warnings.catch_warnings():
|
||||
warnings.simplefilter("error")
|
||||
with pytest.raises(UserWarning) as e:
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
assert str(e.value) == "No tables found on page-1"
|
||||
|
||||
|
||||
def test_stream_no_tables_in_area():
|
||||
filename = os.path.join(testdir, "only_page_number.pdf")
|
||||
with warnings.catch_warnings():
|
||||
warnings.simplefilter("error")
|
||||
with pytest.raises(UserWarning) as e:
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
assert str(e.value) == "No tables found in table area 1"
|
||||
|
||||
|
||||
def test_no_tables_found_logs_suppressed():
|
||||
filename = os.path.join(testdir, 'foo.pdf')
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
with warnings.catch_warnings():
|
||||
# the test should fail if any warning is thrown
|
||||
warnings.simplefilter('error')
|
||||
warnings.simplefilter("error")
|
||||
try:
|
||||
tables = camelot.read_pdf(filename, suppress_stdout=True)
|
||||
except Warning as e:
|
||||
warning_text = str(e)
|
||||
pytest.fail('Unexpected warning: {}'.format(warning_text))
|
||||
pytest.fail(f"Unexpected warning: {warning_text}")
|
||||
|
||||
|
||||
def test_no_tables_found_warnings_suppressed():
|
||||
filename = os.path.join(testdir, 'blank.pdf')
|
||||
filename = os.path.join(testdir, "empty.pdf")
|
||||
with warnings.catch_warnings():
|
||||
# the test should fail if any warning is thrown
|
||||
warnings.simplefilter('error')
|
||||
warnings.simplefilter("error")
|
||||
try:
|
||||
tables = camelot.read_pdf(filename, suppress_stdout=True)
|
||||
except Warning as e:
|
||||
warning_text = str(e)
|
||||
pytest.fail('Unexpected warning: {}'.format(warning_text))
|
||||
pytest.fail(f"Unexpected warning: {warning_text}")
|
||||
|
||||
|
||||
def test_no_password():
|
||||
filename = os.path.join(testdir, 'health_protected.pdf')
|
||||
message = 'file has not been decrypted'
|
||||
filename = os.path.join(testdir, "health_protected.pdf")
|
||||
message = "file has not been decrypted"
|
||||
with pytest.raises(Exception, match=message):
|
||||
tables = camelot.read_pdf(filename)
|
||||
|
||||
|
||||
def test_bad_password():
|
||||
filename = os.path.join(testdir, 'health_protected.pdf')
|
||||
message = 'file has not been decrypted'
|
||||
filename = os.path.join(testdir, "health_protected.pdf")
|
||||
message = "file has not been decrypted"
|
||||
with pytest.raises(Exception, match=message):
|
||||
tables = camelot.read_pdf(filename, password='wrongpass')
|
||||
tables = camelot.read_pdf(filename, password="wrongpass")
|
||||
|
||||
@@ -11,57 +11,50 @@ testdir = os.path.dirname(os.path.abspath(__file__))
|
||||
testdir = os.path.join(testdir, "files")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_text_plot():
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
return camelot.plot(tables[0], kind='text')
|
||||
return camelot.plot(tables[0], kind="text")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_grid_plot():
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
return camelot.plot(tables[0], kind='grid')
|
||||
return camelot.plot(tables[0], kind="grid")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_lattice_contour_plot():
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
return camelot.plot(tables[0], kind='contour')
|
||||
return camelot.plot(tables[0], kind="contour")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_stream_contour_plot():
|
||||
filename = os.path.join(testdir, "tabula/12s0324.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor='stream')
|
||||
return camelot.plot(tables[0], kind='contour')
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
return camelot.plot(tables[0], kind="contour")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_line_plot():
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
return camelot.plot(tables[0], kind='line')
|
||||
return camelot.plot(tables[0], kind="line")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_joint_plot():
|
||||
filename = os.path.join(testdir, "foo.pdf")
|
||||
tables = camelot.read_pdf(filename)
|
||||
return camelot.plot(tables[0], kind='joint')
|
||||
return camelot.plot(tables[0], kind="joint")
|
||||
|
||||
|
||||
@pytest.mark.mpl_image_compare(
|
||||
baseline_dir="files/baseline_plots", remove_text=True)
|
||||
@pytest.mark.mpl_image_compare(baseline_dir="files/baseline_plots", remove_text=True)
|
||||
def test_textedge_plot():
|
||||
filename = os.path.join(testdir, "tabula/12s0324.pdf")
|
||||
tables = camelot.read_pdf(filename, flavor='stream')
|
||||
return camelot.plot(tables[0], kind='textedge')
|
||||
tables = camelot.read_pdf(filename, flavor="stream")
|
||||
return camelot.plot(tables[0], kind="textedge")
|
||||
|
||||