使用python中的Selenium Webdriver下载在embed标签中具有stream-url是chrome扩展名的文件

在Firefox页面中用于<object data="...">显示带有扫描的PDF。“上传的文档”部分中有按钮，以显示其他扫描。

这段代码使用这些按钮来显示扫描，从中获取数据<object>并保存在文件document-0.pdf中 document-1.pdf，等等。

我使用的是您在上一个问题的答案中看到的相同代码：在python中使用Selenium Webdriver保存pdf

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import webdriverwait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.keys import Keys
import time

url = 'https://maharerait.mahaonline.gov.in'

#chrome_path = r'C:/Users/User/AppData/Local/Programs/Python/python36/Scripts/chromedriver.exe'
#driver = webdriver.Chrome(executable_path=chrome_path)

driver = webdriver.Firefox()

driver.get(url)

webdriverwait(driver, 20).until(EC.element_to_be_clickable((By.XPATH,"//div[@class='search-pro-details']//a[contains(.,'Search Project Details')]"))).click()
registered_project_radio = webdriverwait(driver, 10).until(EC.element_to_be_clickable((By.ID,"Promoter")))
driver.execute_script("arguments[0].click();", registered_project_radio)

application = driver.find_element_by_id("CertiNo")
application.send_keys("P50500000005")

search = webdriverwait(driver, 10).until(EC.element_to_be_clickable((By.ID,"btnSearch")))
driver.execute_script("arguments[0].click();", search)

time.sleep(5)

View = [item.get_attribute('href')
         for item in driver.find_elements_by_tag_name("a")
          if item.get_attribute('href') is not None]

# if there is list then get first element
if View:
    View = View[0]

#-----------------------------------------------------------------------------

# load page    
driver.get(View)

# find buttons in section `Uploaded Documents`
buttons = driver.find_elements_by_xpath('//div[@id="DivDocument"]//button')

# work with all buttons 
for i, button in enumerate(buttons):

    # click button
    button.click()

    # wait till page display scan
    print('wait for object:', i)
    search = webdriverwait(driver, 10).until(EC.visibility_of_element_located((By.TAG_NAME, "object")))

    # get data from object    
    print('get data:', i)    
    import base64

    obj = driver.find_element_by_tag_name('object')
    data = obj.get_attribute('data')
    text = data.split(',')[1]
    bytes = base64.b64decode(text)

    # save scan in next PDF     
    print('save: document-{}.pdf'.format(i))    
    with open('document-{}.pdf'.format(i), 'wb') as fp:
        fp.write(bytes)

    # close scan        
    print('close document:', i)    
    driver.find_element_by_xpath('//button[text()="Close"]').click()

# --- end ---

driver.close()

python 2022/1/1 18:32:45 有210人围观

撰写回答

你尚未登录，登录后可以

和开发者交流问题的细节

关注并接收问题和回答的更新提醒

参与内容的编辑和改进，让解决方法与时俱进

请先登录

使用python中的Selenium Webdriver下载在embed标签中具有stream-url是chrome扩展名的文件

撰写回答

推荐问题

Greasemonkey 1.0中的jQuery与使用jQuery的网站冲突

如何使用JSON-LD标记面包屑列表中的最后一个非链接项目

如何在Spring MVC中使用AJAX渲染视图

使用动态where子句休眠

如何使用jQuery访问父窗口对象？

使用Curl和PHP使会话保持活动状态

如何建立一个动态查询，该查询增加了迄今为止的天数，并使用标准API比较该日期与另一个日期？

使用LESS构建选择器列表

如何使用CSS将跨度更改为类似pre？

在mysql sproc中使用变量作为表名

如何使用C＃获取两个DateTime对象之间的时差？

我可以在php中的SESSION数组上使用array_push吗？

Django-如何使用South重命名模型字段？

使用Spring Functional Web Framework的REST端点的背压

使用GhostDriver时如何设置屏幕/窗口大小

如何使用最新版本的jQuery并在RichFaces中为jQuery取回“ $”？

我可以使用BeautifulSoup删除脚本标签吗？

多态对象的JSON使用者

我如何重新连接使用selenium的webdriver打开的浏览器？

如何使用Servlet和Ajax？

分类汇总

您的鼓励是对我最大的支持