I'm using nokogiri to parse an XML file. Some of the nodes in the file have attributes specific to namespaces:
<metadata xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:opf="http://www.idpf.org/2007/opf">
<dc:identifier id="iden" opf:scheme="ISBN">xxxx</dc:identifier>
<dc:creator opf:role="aut" opf:file-as="Name">xxxx</dc:creator>
<dc:date opf:event="publication">xxxx</dc:date>
<dc:publisher>xxxx</dc:publisher>
<meta name="cover" content="x"/>
</metadata>
I'm trying to remove any attribute with the "opf" prefix. I've come across xpath solutions in finding an attribute value based on a partial match, but what about when it's a partial match of the attribute name itself? I tried a lot of things that haven't worked. I did a simple thing just to try to extract the attribute names at least, but if I do:
elements = @doc.at_xpath('//xmlns:metadata').children
elements.each { |el|
el.attributes.each { |attribute|
if attribute[1].namespace_scopes[1].prefix == "opf"
puts attribute[0]
end
}
}
I end up getting:
id
scheme
role
file-as
event
name
content
but I only want the ones with the "opf" prefix ("opf:scheme", "opf:role, "opf:file-as", "opf:event") so that they can be removed, without touching any of the other attributes. I even tried to force it by hard-coding the attributes I knew existed:
opf_attributes = ["opf:file-as","opf:scheme","opf:role","opf:event"]
elements.each { |el|
opf_attributes.each { |x|
el.remove_attribute(x) if el[x] != nil
}
}
which is not the smartest way to go about this, but this still didn't work. Nothing happens to the nodes, and the attributes remain as they were. (I don't know if it's worth noting, but if I use the remove_attr(x) method instead, I get this error: undefined method 'remove_attr' for #<Nokogiri::XML::Element:0x...>
So, my question is:
Is there a clearer way to
- find attributes based on a partial match and/or the namespace prefix, then
- remove those attributes from the nodes that contain them?