提问人:Rodrigo 提问时间:6/28/2021 最后编辑:Rodrigo 更新时间:7/17/2021 访问量:5293
硒打印 A4 格式的 PDF
Selenium print PDF in A4 format
问:
我有以下代码用于打印到 PDF(并且它可以工作),并且我只使用 Google Chrome 进行打印。
def send_devtools(driver, command, params=None):
# pylint: disable=protected-access
if params is None:
params = {}
resource = "/session/%s/chromium/send_command_and_get_result" % driver.session_id
url = driver.command_executor._url + resource
body = json.dumps({"cmd": command, "params": params})
resp = driver.command_executor._request("POST", url, body)
return resp.get("value")
def export_pdf(driver):
command = "Page.printToPDF"
params = {"format": "A4"}
result = send_devtools(driver, command, params)
data = result.get("data")
return data
正如我们所看到的,我用于打印到 base64,并像在 paramenter 上一样传递“A4”。Page.printToPDF
format
params
不幸的是,这个参数似乎被忽略了。我看到一些使用puppeteer的代码(格式A4),我认为这可以帮助我。
即使有硬编码的宽度和高度(见下文),我也没有运气。
"paperWidth": 8.27, # inches
"paperHeight": 11.69, # inches
使用上面的代码,是否可以将页面设置为 A4 格式?
答:
更新帖子 07-17-2021
我决定使用 Python 包 pdfminer.sixth 验证我的原始代码的输出
from pdfminer.pdfpage import PDFPage
from pdfminer.pdfpage import PDFParser
from pdfminer.pdfpage import PDFDocument
parser = PDFParser(open('test_1.pdf', 'rb'))
doc = PDFDocument(parser)
pageSizesList = []
for page in PDFPage.create_pages(doc):
print(page.mediabox)
# output
[0, 0, 612, 792]
当我将这些点大小转换为英寸时,我感到震惊。尺寸为 8.5 x 11,不等于 A4 纸张尺寸 8.27 x 11.69。当我看到这一点时,我决定通过查看 chromium 和 selenium 源代码来进一步探索这个问题。
在 chromium 源代码中,该命令位于文件 page_handler.cc 中Page.printToPDF
void PageHandler::PrintToPDF(Maybe<bool> landscape,
Maybe<bool> display_header_footer,
Maybe<bool> print_background,
Maybe<double> scale,
Maybe<double> paper_width,
Maybe<double> paper_height,
Maybe<double> margin_top,
Maybe<double> margin_bottom,
Maybe<double> margin_left,
Maybe<double> margin_right,
Maybe<String> page_ranges,
Maybe<bool> ignore_invalid_page_ranges,
Maybe<String> header_template,
Maybe<String> footer_template,
Maybe<bool> prefer_css_page_size,
Maybe<String> transfer_mode,
std::unique_ptr<PrintToPDFCallback> callback)
此功能允许修改参数和。这些参数采用 A C++ double 是一种通用数据类型,编译器在内部用于定义和保存任何数值数据类型,尤其是任何面向十进制的值。C++ double 数据类型可以是小数,也可以是带有值的整数。paper_width
paper_height
double.
这些参数具有默认值,这些值在 Chrome DevTools 协议中定义:
- paperWidth:纸张宽度(以英寸为单位)。默认为 8.5 英寸。
- paperHeight:纸张高度(以英寸为单位)。默认为 11 英寸
请注意源代码和详细信息之间的参数格式之间的差异。chromium
Chrome DevTools Protocol
- 源代码中的paper_width
chromium
- paperWidth 中的
Chrome DevTools Protocol
根据源代码,该命令被调用chromium
Page.printToPDF
SendCommandAndGetResultWithTimeout.
Status WebViewImpl::PrintToPDF(const base::DictionaryValue& params,
std::string* pdf) {
// https://bugs.chromium.org/p/chromedriver/issues/detail?id=3517
if (!browser_info_->is_headless) {
return Status(kUnknownError,
"PrintToPDF is only supported in headless mode");
}
std::unique_ptr<base::DictionaryValue> result;
Timeout timeout(base::TimeDelta::FromSeconds(10));
Status status = client_->SendCommandAndGetResultWithTimeout(
"Page.printToPDF", params, &timeout, &result);
if (status.IsError()) {
if (status.code() == kUnknownError) {
return Status(kInvalidArgument, status);
}
return status;
}
if (!result->GetString("data", pdf))
return Status(kUnknownError, "expected string 'data' in response");
return Status(kOk);
}
在我最初的答案中,我使用了
与命令类似send_command_and_get_result,
SendCommandAndGetResultWithTimeout.
# stub_devtools_client.h
Status SendCommandAndGetResult(
const std::string& method,
const base::DictionaryValue& params,
std::unique_ptr<base::DictionaryValue>* result) override;
Status SendCommandAndGetResultWithTimeout(
const std::string& method,
const base::DictionaryValue& params,
const Timeout* timeout,
std::unique_ptr<base::DictionaryValue>* result) override;
在查看了 selenium 源代码后,不清楚如何正确传递命令或send_command_and_get_result
send_command_and_get_result_with_timeout.
我确实在 selenium 源代码中注意到了这个函数:webdriver
def execute_cdp_cmd(self, cmd, cmd_args):
"""
Execute Chrome Devtools Protocol command and get returned result
The command and command args should follow chrome devtools protocol domains/commands, refer to link
https://chromedevtools.github.io/devtools-protocol/
:Args:
- cmd: A str, command name
- cmd_args: A dict, command args. empty dict {} if there is no command args
:Usage:
driver.execute_cdp_cmd('Network.getResponseBody', {'requestId': requestId})
:Returns:
A dict, empty dict {} if there is no result to return.
For example to getResponseBody:
{'base64Encoded': False, 'body': 'response body string'}
"""
return self.execute("executeCdpCommand", {'cmd': cmd, 'params': cmd_args})['value']
经过一些研究和测试,我发现这个功能可以用来实现你的用例。
import base64
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from pdfminer.pdfpage import PDFPage
from pdfminer.pdfpage import PDFParser
from pdfminer.pdfpage import PDFDocument
chrome_options = Options()
chrome_options.add_argument("--disable-infobars")
chrome_options.add_argument("--disable-extensions")
chrome_options.add_argument("--disable-popup-blocking")
chrome_options.add_argument('--headless')
browser = webdriver.Chrome('/usr/local/bin/chromedriver', options=chrome_options)
browser.get('http://www.google.com')
# use can defined additional parameters if needed
params = {'landscape': False,
'paperWidth': 8.27,
'paperHeight': 11.69}
# call the function "execute_cdp_cmd" with the command "Page.printToPDF" with
# parameters defined above
data = browser.execute_cdp_cmd("Page.printToPDF", params)
# save the output to a file.
with open('file_name.pdf', 'wb') as file:
file.write(base64.b64decode(data['data']))
browser.quit()
# verify the page size of the PDF file created
parser = PDFParser(open('file_name.pdf', 'rb'))
doc = PDFDocument(parser)
pageSizesList = []
for page in PDFPage.create_pages(doc):
print(page.mediabox)
# output
[0, 0, 594.95996, 840.95996]
输出以磅为单位,需要转换为英寸。
- 594.95996 点 等于 8.263332777783 英寸
- 840.95996 积分 等于 11.6799994445英寸
8.263332777783 x 11.6799994445 是 A4 纸尺寸。
原帖 07-13-2021
调用函数时可以传递多个参数 其中两个参数是:Page.printToPDF.
- paper_width
- paper_height
以下代码将这些参数传递给Page.printToPDF.
import json
import base64
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
def send_devtools(driver, command, params=None):
if params is None:
params = {}
resource = "/session/%s/chromium/send_command_and_get_result" % driver.session_id
url = driver.command_executor._url + resource
body = json.dumps({"cmd": command, "params": params})
resp = driver.command_executor._request("POST", url, body)
return resp.get("value")
def create_pdf(driver, file_name):
command = "Page.printToPDF"
params = {'paper_width': '8.27', 'paper_height': '11.69'}
result = send_devtools(driver, command, params)
save_pdf(result, file_name)
return
def save_pdf(data, file_name):
with open(file_name, 'wb') as file:
file.write(base64.b64decode(data['data']))
print('PDF created')
chrome_options = Options()
chrome_options.add_argument("--disable-infobars")
chrome_options.add_argument("--disable-extensions")
chrome_options.add_argument("--disable-popup-blocking")
chrome_options.add_argument('--headless')
browser = webdriver.Chrome('/usr/local/bin/chromedriver', options=chrome_options)
browser.get('http://www.google.com')
create_pdf(browser, 'test_pdf_1.pdf')
----------------------------------------
My system information
----------------------------------------
Platform: maxOS
OS Version: 10.15.7
Python Version: 3.9
Selenium: 3.141.0
pdfminer.sixth: 20201018
----------------------------------------
评论