# python - wget xml doc and parse with awk

**URL:** <https://community.unix.com/t/python-wget-xml-doc-and-parse-with-awk/295236>\
**Category:** Shell Programming and Scripting\
**Created:** [September 14, 2011, 2:10am UTC](https://community.unix.com/t/python-wget-xml-doc-and-parse-with-awk/295236 "2011-09-14T02:10:42Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![unclecameron](https://community.unix.com/letter_avatar/unclecameron/32/5_5575768a8748004e209b776fc1b2916d.png) [@unclecameron](https://community.unix.com/u/unclecameron)\
**Post date:** [September 14, 2011, 2:10am UTC](https://community.unix.com/t/python-wget-xml-doc-and-parse-with-awk/295236/1 "2011-09-14T02:10:42Z")

</div>

Well, that's what I'd do in bash 🙂 Here's what I have so far:

```nohighlight
import urllib2
from BeautifulSoup import BeautifulStoneSoup

xml = urllib2.urlopen('http://weatherlink.com/xml.php?user=blah&pass=blah')
soup = BeautifulStoneSoup(xml)
print soup.prettify()

```

but all it does is grab the html page, not the xml it should serve. The example xml looks like:

```nohighlight
...
<title>blah</title>
<link>http://www.blah.com</link>
</image>
<suggested_pickup>15 minutes after the hour</suggested_pickup>
<dewpoint_c>16.7</dewpoint_c>
<dewpoint_f>62</dewpoint_f>
<heat_index_f>77</heat_index_f>
...

```

what can I do to make:

```nohighlight
some_data {}
some_data ['dewpoint_c'] = 16.7
some_data ['heat_index'] = 77

```

or whatever those values should be from the xml it should get? I've also tried things like:

```nohighlight
dom = minidom.parse(urllib2.urlopen(url))

&

xml = ElementTree.fromstring(result)
print xml.findtext(".//heat_index_f")

```

but don't seem to be getting anywhere.
