推荐学习书目
› Learn Python the Hard Way
Python Sites
› PyPI - Python Package Index
› http://diveintopython.org/toc/index.html
› Pocoo
值得关注的项目
› PyPy
› Celery
› Jinja2
› Read the Docs
› gevent
› pyenv
› virtualenv
› Stackless Python
› Beautiful Soup
› 结巴中文分词
› Green Unicorn
› Sentry
› Shovel
› Pyflakes
› pytest
Python 编程
› pep8 Checker
Styles
› PEP 8
› Google Python Style Guide
› Code Style from The Hitchhiker's Guide
tioover
V2EX  ›  Python

python 如何过滤 HTML标签?

  •  
  •   tioover ·
    tioover · May 21, 2012 · 14050 views
    This topic created in 5243 days ago, the information mentioned may be changed or developed.
    17 replies  •  1970-01-01 08:00:00 +08:00
    wong2
        1
    wong2  
       May 21, 2012   ❤️ 1
    VeryCB
        2
    VeryCB  
       May 21, 2012   ❤️ 1
    VeryCB
        3
    VeryCB  
       May 21, 2012
    lackrp
        4
    lackrp  
       May 21, 2012   ❤️ 1
    过滤是指要去掉么?

    import re
    pattern = re.compile(r'<.*?>')
    pattern.sub('', html)
    gee
        5
    gee  
       May 21, 2012
    @VeryCB 你说得一点都不靠谱吧...你给的都是解析html的...
    phuslu
        6
    phuslu  
       May 21, 2012   ❤️ 1
    python readability
    VeryCB
        7
    VeryCB  
       May 21, 2012
    @gee 初学者...过滤不是把Html标签去掉然后提取内容么?
    eric_q
        8
    eric_q  
       May 21, 2012
    我……我还是想用shell
    gee
        9
    gee  
       May 21, 2012
    @VeryCB 问题是不了解具体的html结构啊。当然了,用PyQuery直接取全体的text()也可以,但是有点绕路了
    TheC
        10
    TheC  
       May 21, 2012
    @lackrp 你这哪里是过滤html标签,误伤也太大了
    lackrp
        11
    lackrp  
       May 21, 2012
    @TheC 正确的html的话,这样应该不会误伤啊
    eerie
        12
    eerie  
       May 21, 2012   ❤️ 1
    TheC
        13
    TheC  
       May 21, 2012
    @lackrp 仔细想想确实是,抱歉刚才回的太快了,理所当然觉得除了html标签还有其他被<>包括的文本了:)
    tioover
        14
    tioover  
    OP
       May 21, 2012
    @lackrp 额。我的意思是过滤掉诸如<sript><iframe>之类的标签只允许一些安全的标签,为了安全……
    cute
        15
    cute  
       May 22, 2012
    @tioover

    from BeautifulSoup import BeautifulSoup
    soup = BeautifulSoup('<html><p>abc</p><script></script><br />abc</html>')
    for tag in soup.recursiveChildGenerator():
    .... if hasattr(tag, 'name') and tag.name in ['script', 'iframe']:
    ........tag.extract()
    print soup.renderContents('utf-8')
    leiz
        16
    leiz  
       May 22, 2012
    @tioover

    如果只是过滤部分标签,干嘛不直接穷举好了?
    magicshui
        17
    magicshui  
       May 22, 2012
    BeautifulSoup感觉简单些~
    About   ·   Help   ·   Advertise   ·   Blog   ·   API   ·   FAQ   ·   Privacy   ·   Solana   ·   5423 Online   Highest 6679   ·     Select Language
    创意工作者们的社区
    World is powered by solitude
    VERSION: 3.9.8.5 · 88ms · UTC 02:59 · PVG 10:59 · LAX 19:59 · JFK 22:59
    ♥ Do have faith in what you're doing.