[
跳过导航]
≡
↩️
🗣️
-
🏠
Help
:
📑
🏁
:
CMS Detectors
≡
欢迎。
N.签名
创建帐户
CMS Detectors@Help
看
来源
历史
讨论
Help组
创建/查找页面
小组讨论
我的团体
🕰️
区域设置:en-US
页:CMS Detectors
⚙
🗄️
页类型:
标准
页和反馈
页别名
媒体名单
演示
Url缩短器
共享墙
Git Repository
Front Page
News Article
别网页为:
页边境:
固体
虚
没有
表中的内容:
标题:
作者:
Meta机器人:
Meta Description:
元属性([2][8]),如开放图([2][9])
格式为:名称|内容的每个属性一行
标题页名称:
脚页名称:
'''CMS Detectors''' are used to help Yioop get to the most important content on a web page. <br /><br /> You must enter the '''Name'''. The Header Regex and Important Content XPath are optional but will have no effect if they are not entered. <br /> '''The Header Regex''' is used to detect the CMS. The header of most CMS created sites are very common. A specifically crafted regular expression can be used to detect the CMS you are looking for. It looks in the href value in a rel='stylesheet' tag or the src value in a type='text/javascript' tag. <br /><br /> The '''Important Content XPath''' is used to target the most important content for summarizing. The first entry is where to target the important content. Any subsequent entry will be used to remove content within the important content. Append each removal XPath to the end of the value delimited by three pound signs (###). <br /> '''Example:''' <br /><br /> <table border='1'> <th>Setting</th> <th>Value</th> <tr><td>Name</td><td>Wordpress</td></tr> <tr><td>Header Regex </td><td>wp-(?:content|includes)</td></tr> <tr><td>Important Content XPath</td><td>//div[@id="content"]###<br />//div[@id="comments"]###<br />//div[@id="respond"]</td></tr> </table> <br />
X
(c)这个网站 -
这个搜索引擎