Web
Description
The NodeBox Web library offers a collection of services to retrieve content from the internet. You can use the library to query Yahoo! for links, images, news and spelling suggestions, to read RSS and Atom newsfeeds, to retrieve articles from Wikipedia, to collect quality images from morgueFile or Flickr, to get color themes from kuler or Colr, to browse through HTML documents, to clean up HTML, to validate URL's, to create GIF images from math equations using mimeTeX, to get ironic word definitions from Urban Dictionary.
The NodeBox Web library works with a caching mechanism that stores things you download from the web, so they can be retrieved faster the next time. Many of the services also work asynchronously. This means you can use the library in an animation that keeps on running while new content is downloaded in the background.
The library bundles Leonard Richardson's Beautiful Soup to parse HTM, Mark Pilgrim's Universal Feed Parser for newsfeeds, a connection to John Forkosh's mathTeX server (thanks Cedric Foellmi), Leif K-Brooks entity replace algorithm, simplejson, and patches for Debian from the people at Indywiki.
Download
![]() | web.zip (390KB) Last updated for NodeBox 1.9.4.6 Licensed under GPL Author: Tom De Smedt |
Documentation
- How to get the library up and running
- Validating web content
- Working with URL's
- Working with HTML
- Querying Yahoo! for links, images and news
- Improving Yahoo! results with a contextual search
- Using Yahoo! to suggest spelling corrections
- Using Yahoo! to sort associatively
- Querying Google
- Reading newsfeeds
- Retrieving articles from Wikipedia
- Some helper commands to draw Wikipedia content in NodeBox
- Querying morgueFile for images
- Querying Flickr for images
- Querying kuler for color themes
- Querying Colr for color schemes
- Creating GIF images from math equations
- Word definitions from Urban Dictionary
- Working with asynchronous downloads
- Clearing the cache
- Reading JSON
How to get the library up and running
Put the web library folder in the same folder as your script so NodeBox can find the library.
You can also put it in ~/Library/Application Support/NodeBox/.
web = ximport("web")
Outside of NodeBox you can also just do import web.
Proxy servers
If you are behind a proxy server the library may not be able to connect to the internet.
In that case you need to inform the library with the set_proxy() command:
web.set_proxy("https://www.myproxyserver.com:80", type="https")
Validating web content
Web content is accessed with a URL, the address you use to connect to a place on the internet. The library has a number of commands to find out what type of content (e.g. web page, image, ...) is associated with a given URL.
The most basic command, is_url() checks whether a given string is a grammatically correct URL (e.g. http://nodebox.net but not htp://nodebox.net). It takes a wait parameter indicating the number of seconds after which the library should stop connecting to the internet and give up.
web.is_url(url, wait=10)
Even if a URL is valid, it might not refer to actual content on the internet. We can check if a URL exists with the not_found() command:
web.url.not_found(url, wait=10)
The following commands are useful to find out what content is associated with the URL. We can discern between HTML web pages which we can parse with page.parse(), newsfeeds which we can parse with newsfeed.parse(), images, audio and video etc. which we can download with url.retrieve().
web.url.is_webpage(url, wait=10)
web.url.is_stylesheet(url, wait=10)
web.url.is_plaintext(url, wait=10)
web.url.is_pdf(url, wait=10)
web.url.is_newsfeed(url, wait=10)
web.url.is_image(url, wait=10)
web.url.is_audio(url, wait=10)
web.url.is_video(url, wait=10)
web.url.is_archive(url, wait=10)
Working with URL's
A URL is the address you use to connect to a page on the internet, for example: http://nodebox.net. The NodeBox Web library can do three different things with a URL: download the content associated with it, parse it (find out which parts make up the URL) and construct it from scratch (a simple way to create a URL with HTTP GET or HTTP POST data).
web.download(url, wait=60, cache=None, type=".html")
web.save(url, path="", wait=60)
web.url.retrieve(url, wait=60, asynchronous=False, cache=None, type=".html")
web.url.parse(url)
web.url.create(url="", method="get")
The download() command returns the content associated with the given web address. The command has an optional parameter wait that determines how long to wait for a download. If the time is exceeded, the download is aborted.
The last two parameters can be used to cache downloaded content locally, so it doesn't have to downloaded again in the future. The cache parameter is a string with the name of a subfolder in /cache where to store content. The type parameter is the file extension of the downloaded content.
The save() command stores the URL's content at the given local path. If no path is given it will attempt to extract a filename from the URL and store that in the current working directory. The path to the saved file is returned.
The Web library also has easier ways to deal with specific web content like HTML (page.parse() command) or Wikipedia articles (wikipedia.search() command) for example.
The download() command is actually an alias of the url.retrieve() command. This command has an additional asynchronous parameter useful to download stuff in the background while an animation keeps on running. We'll see about asynchronous downloads later on. The url.retrieve() command returns an object with a data property containing a string with the downloaded content. If you don't need anything that complicated just use the easy download() command:
# Download an image from the NodeBox Gallery. url = "http://nodebox.net/code/data/media/twisted-final.jpg" img = web.download(url) # Display the image data in NodeBox. image(None, 0, 0, data=img) # Write the image data to a file. file = open("twisted.jpg", "w") file.write(img) file.close()
The url.parse() command splits a given url into its components. The returned objects has the following properties:
- url.protocol: the type of internet service, usually http
- url.domain: the domain name, for example, nodebox.net
- url.username: a username for a secure connection
- url.password: a password for a secure connection
- url.port: the port number at the host
- url.path: the subdirectory at the server, for example /code/index.php/
- url.page: the name of the document, for example search
- url.anchor: named anchor on the page
- url.query: a dictionary of query string values, for example { "q": "pixels" }
- url.method: the query string method either "get" or "post"
In the same way the url.create() command returns an object with these properties. This command is useful to, for example, construct URL's with a POST query and pass that to url.retrieve() or page.parse().
For example. this script retrieves the first 10 forum pages from NodeBox:
url = web.url

