Showing posts with label search. Show all posts
Showing posts with label search. Show all posts

Friday, January 20, 2017

SiteTruth, interesting startup

By Vasudev Ram

SiteTruth is an early-stage startup with an interesting product, that I got to know about today.

Excerpt from their About page above:

[ Conventional wisdom is that search spam can't be stopped.
Conventional wisdom is wrong.

SiteTruth exists to solve one of the Web's biggest problems - web spam, unidentified, and possibly fake, on-line businesses.

Every on-line commerce web site must display the name and address of the business behind the site. That's the law in much of the developed world. SiteTruth tries to identify that business, then find information about it. That check is used to influence search rankings. That's SiteTruth.

Sites which sell, but have no identifiable business behind them, are down-rated. "Doorway pages", "affiliate sites", and anonymous businesses receive low ratings. Identifiable, verifiable businesses get good ratings. The marginal businesses, the ones that the customer can't find when they need a refund, get low ratings. ]

- Vasudev Ram - Online Python training and consulting

Get updates (via Gumroad) on my forthcoming apps and content.

Jump to posts: Python * DLang * xtopdf

Subscribe to my blog by email

My ActiveState Code recipes

Follow me on: LinkedIn * Twitter

Managed WordPress Hosting by FlyWheel



Saturday, April 20, 2013

Samuru, a new (and smarter?) search engine


By Vasudev Ram - Dancing Bison Enterprises

Samuru is a new search engine from stremor.com

I tried it out a bit and it gave me some interesting results.

About stremor

About Samuru:

Update in a while, I'm on my mobile at present.

Hacker News thread about Samuru is interesting


- Vasudev Ram

Saturday, April 13, 2013

The pip search command (Python)


By Vasudev Ram

I was checking out the help for Python's pip tool (an installation tool for Python packages), and came across its search option.

The search option of the pip command lets you search PyPI, the Python Package Index.

I tried it out with a couple of searches, for the terms bottle (to get info on more PyPI packages related to 1) the Bottle Python web microframework (which I had blogged about a while ago), and 2) PDF, which is one of my areas of interest.

Here are the results of those two pip searches of PyPI:

c:\Python27\Lib\site-packages>pip -v -v -v search bottle
bottle                    - Fast and simple WSGI-framework for small web-
                            applications.
  INSTALLED: 0.11.6 (latest)
bottlenose                - A Python hook into the Amazon.com Product
                            Advertising API
Bottleneck                - Fast NumPy array functions written in Cython
mimerender                - RESTful HTTP Content Negotiation for Flask,
                            Bottle, web.py and webapp2 (Google App Engine)
bottle-cork               - Authentication/Authorization library for Bottle
pasttle                   - Simple pastebin on top of bottle.
bottle-websocket          - WebSockets for bottle
bottle-hotqueue           - FIFO Queue for Bottle built upon HotQueue
Bottle-DebugToolbar       - A port of the Django Debug Toolbar to Bottle
bottle-haml               - UNKNOWN
macaron                   - Simple object-relational mapper for SQLite3,
                            includes plugin for Bottle web framework
bottle-servefiles         - A reusable app that serves static files for bottle
                            apps
bottle-sqlalchemy         - SQLAlchemy integration for Bottle.
bottle-mongodb            - MongoDB integration for Bottle
flaskle                   - bottle-like utility decorators for flask
bottle-redis              - Redis integration for Bottle.
bottle-pystache           - Bottle Pystache template wrappers
bottle-tornado-websocket  - WebSockets for bottle
BottleRack                - BottleRack Markdown HTML Server
Gluino                    - port of web2py libs to bottle, flask, pyramid,
                            tornado (includes copy of modules from the web2py
                            framework)
bottle-request            - Plugin to give bottle a 'stateless' request object
bottle-sqlite             - SQLite3 integration for Bottle.
bottle-renderer           - Renderer plugin for bottle
bottle-memcache           - Memcache integration for Bottle.
bottle-web2pydal          - Web2py Dal integration for Bottle.
Bottle-SSLify             - Force SSL on any Bottle app.
bottle-werkzeug           - Werkzeug integration for Bottle.
bottle-pycassa            - Bottle plugin for Cassandra/Pycassa
bottle_cql                - CQL integration for Bottle.
bottle-extras             - Meta package to install the bottle plugin
                            collection.
bn                        - Lightweight profiling tool to detect performance
                            BottleNecks in Python code.
bottle-agamemnon          - Agamemnon integration for bottle
bottle-flash              - flash plugin for bottle
bottle-tornadosocket      - WebSockets for bottle
bottle-mysql              - MySQL integration for Bottle.
bottle-pgsql              - PgSQL integration for Bottle

c:\Python27\Lib\site-packages>
c:\Python27\Lib\site-packages>pip -v -v -v search pdf
mwlib.rl                  - generate pdfs from mediawiki markup
zopyx.convert             - A Python interface to XSL-FO libraries (Conversion
                            HTML to PDF, RTF, DOCX, WML and ODT)
slc.publications          - A content type to store and parse pdf publications
pdfminer                  - PDF parser and analyzer
zopyx.convert2            - A Python interface for the conversion of HTML to
                            PDF, RTF, DOCX, WML and ODT) - belongs to
                            zopyx.smartprintng.core
pisa                      - PDF generator using HTML and CSS
collective.pdfpeek        - A Plone 4 product that generates image thumbnail
                            previews of PDF files stored on ATFile based
                            objects.
WeasyPrint                - WeasyPrint converts web documents to PDF.
rst2pdf                   - Convert restructured text to PDF via reportlab.
pdfserenitynow            - Create TIFs and JPGs from crappy PDFs
collective.sendaspdf      - An open source product for Plone to download or
                            email a page seen by the user as a PDF file.
Products.SmartPrintNG     - Produce & Publish for Plone - Conversion of Plone
                            content to PDF, RTS, ODT, DOCX and WML
wc.pageturner             - A Plone product that provides the PDF viewer
                            FlexPaper.
PyX                       - Python package for the generation of PostScript
                            and PDF files
PollyReports              - Band-oriented PDF report generation from database
                            query
relatorio                 - A templating library able to output odt and pdf
                            files
pyPdf                     - PDF toolkit
wildcard.pdfpal           - PDF Thumbnail generation, OCR indexing and extra
                            views integrated with plone.app.async
pdfminer3k                - PDF parser and analyzer
pdfparanoia               - pdf watermark remover library for academic papers
template2pdf              - Renders Django/Jinja2 templates into PDF.
ftw.book                  - Produce books with Plone and export them in a high
                            quality PDF.
collective.pdfjs          - pdf.js integration for Plone
ftw.pdfgenerator          - A library for generating PDF representations of
                            Plone objects with LaTeX.
django-wkhtmltopdf        - Converts html to PDF using
                            http://code.google.com/p/wkhtmltopdf/.
TableFactory              - Easily create HTML, spreadsheet, or PDF tables
                            from common Python data sources
paPyro                    - A PDF report generator written in Python
pdfnup                    - Layout multiple pages per sheet of a PDF document.
pdfdocument               - Wrapper for ReportLab which allows easy creation
                            of PDF documents.
xhtml2pdf                 - PDF generator using HTML and CSS
ServPDF                   - ServPDF is a webbased Microsoft Office to PDF
                            Converter
pdfcrowd                  - A client for Pdfcrowd API.
pdfrecycle                - create a PDF file by composing pages from other
                            PDF files.
metapdf                   - A lightweight PDF library optimized for metadata
                            extraction and insertion
tex                       - Convert LaTeX or TeX source to PDF or DVI, and
                            escape strings for LaTeX.
pdfkit                    - Wkhtmltopdf python wrapper to convert html to pdf
                            using the webkit rendering engine and qt
eea.converter             - SVG, PNG, PDF converters using external tools as
                            ImageMagick
pieberry-library-assistant - A program to download pdf documents from public
                            websites, and catalogue them in BibTeX format
InvoiceGenerator          - Library to generate PDF invoice.
pdfserver                 - Pdfserver is a webservice that offers common PDF
                            operations like joining documents, selecting pages
                            or "n pages on one".
rputils                   - An application for dice rolling and reading PDF
                            files.
diffpy.pdfgui             - GUI for PDF simulation and structure refinement.
pyf.components.consumers.rmlpdfwriter - PyF component RML-Based PDF Writer based
 on
                            Z3C.RML and Genshi
django-xhtml2pdf          - A Django app to generate pdfs from templates
aws.pdfbook               - Download Plone content views as PDF
pdfquery                  - Concise and friendly PDF scraper using JQuery or
                            XPath selectors.
OPAF                      - Open PDF Analysis Framework
TecUtils                  - Various utilities for database and config files
                            use. Text to Pdf converter
Products.PDFtoOCR         - PDFtoOCR does OCR processing on PDF documents. The
                            text from OCR is used in the search results.
Flask-WeasyPrint          - Make PDF in your Flask app with WeasyPrint.
pdfmerge                  - Command-line PDF utility.
ftwbook.graphicblock      - Addon for `ftw.book` providing a graphics block
                            for including PDF documents in the book.
pyf.components.consumers.ooowriter - py3o-powered OpenOffice.org ODT Writer for
PyF
                            Framework with support for rendering (pdf, html,
                            doc, docx, etc) on py3o renderservers.
pyjon.reports             - Pyjon.Reports is a module bridging z3c.rml, genshi
                            and pypdf together to provide a simple mean of
                            creating templated pdf documents in python.
django-pdf                - A Django app for managing and processing PDF
                            documents.
gametex-django-print      - Generate PDFs from GameTeX in Django
collective.pdftransform   - A set of portal transform to change pdf into
                            images
pdftools.pdfposter        - Scale and tile PDF images/pages to print on
                            multiple pages.
django-webodt             - ODF template handler and odt to html, pdf, doc,
                            etc converter
slate                     - Extract text from PDF documents easily.
dinbrief                  - PDF renderer for DIN 5008 and DIN 676 compliant
                            letters and invoices
raptus.princexml          - Provides a simple bridge to create PDFs from HTML
                            views using PrinceXML
wallaby-plugin-pdfgenerator - This package provides a PDF generator for wallaby.

z3c.pdftemplate           - PDF Template
wkhtmltopdf               - Simple python wrapper for wkhtmltopdf
diffpy.pdffit2            - PDFfit2 - real space structure refinement program.
pdfid_PL                  - A Python module to analyze and sanitize PDF files,
                            based on Didier Stevens' PDFiD
pypdflib                  - Pango Cairo based Python PDF Library
pdf2zip                   - pdf conversion utility
pyf.components.consumers.xhtmlpdfwriter - XHTML-Based PDF Writer for PyF Framewo
rk based on
                            PISA and Genshi
pdf-link-checker          - Reports broken hyperlinks in PDF documents
scrape-highlighted        - A script to scrape highlighted text from a pdf
                            [Mac only].
pdfgrid                   - Add a grid on top of all pages of a PDF document.
jag-tipdf                 - Combines plain text and images into a single PDF.
HWFormatter               - Format submitted homework assignments from
                            Blackboard into PDFs
collective.pdfLeadImage   - Automatically creates contentleadimage from pdf
                            cover
pdfcat                    - pdf concatenation tool
prynt                     - Generating HTML/PDF directly from your console.
pdfsplit                  - Split a PDF file or rearrange its pages into a new
                            PDF file.
pdfrw                     - Pure Python PDF file reader/writer library
agenda2pdf                - Simple script which generates a book agenda file
                            in PDF format, ready to be printed or to be loaded
                            on a ebook reader
stapler                   - Manipulate PDF documents from the command line
trml2pdf                  - Tiny RML2PDF is a tool to easily create PDF
                            document without programming. It       can be used
                            as a Python library or as a standalone binary. It
                            converts a RML,       an XML dialect that lets you
                            define the precise appearance of a printed
                            document, to a PDF. You can use your existing
                            tools to generate an input file       that exactly
                            describes the layout of a printed document, and
                            RML2PDF converts       it into PDF. RML is a much
                            more powerfull and flexible alternative to XSL:FO.
                            The executable read a RML file to the standard
                            input and output a PDF file to       the standard
                            output.
wsgitrml2pdf              - wsgitrml2pdf is wsgi middleware to convert trml
                            text to pdf.
fpdf.py                   - Simple PDF generation for Python (FPDF PHP port)
DXF-Converter             - Converts Autocad Files v11 and v12 to pdf
JagPDF                    - A library for generating PDF documents.
pdftools.pdfjoin          - Join PDF documents into a single document.
collective.calameo        - Publish your PDF document in Plone with Calameo
PyPDF2                    - PDF toolkit
PDFTron PDFNet SDK for Python - A top notch PDF library for PDF rendering,
                            conversion, content extraction, etc
cubicweb-pdfexport        - export any page as a pdf document
fpdf                      - Simple PDF generation for Python
pywkher                   - wkhtmltopdf for Python on Heroku
PyBooklet                 - Converts PDFs to booklets
pyFPDF                    - generate pdf on python
origapy                   - A Python module to clean PDF files by disabling
                            active content (javascript, launch, etc), using
                            the Ruby Origami PDF parser.
pdftable                  - pdftable: extract tables from PDF files
django-latex              - Django application to generate latex/pdf files.
ttfpdf                    - TrueType support PDF file Generator
pdf2tiff                  - A PDF to TIFF converter for Mac OS X.
cloudooo.handler.pdf      - Python Package to handler PDF documents
pdfbox                    - JCC wrapper for Apache PDFBox
buzzweb2pdf               - An Open Source tool to convert HTML documentation
                            with an index page into a single PDF.
ServPDF and ServPDF-OO    - Web based Office to PDF Converter Server for
                            Microsoft Office and OpenOffice
genenga                   - Generate Nengajo(Japanese new year card) pdf from
                            address list.

As you can see from the output above, a lot of Python packages were found for those two search terms. In the search for the term "bottle", not all of the packages found were actually related to the Bottle framework, but most were.

The search for "PDF" resulted in many results, some of which I've blogged about in the past, such as pisa/xhtmltopdf, PyPDF, pdfdocument, wkhtml2pdf, and others, and many that I didn't know of before.

So it looks like the pip search command is a useful tool.

As for why I used "pip -v -v -v", see "pip --help" :-)

- Vasudev Ram - Dancing Bison Enterprises

Friday, January 11, 2013

pyelasticsearch and elasticsearch, distributed search based on Lucene


pyelasticsearch 0.3 : Python Package Index

pyelasticsearch is a Python wrapper for the elasticsearch distributed search engine (based on Apache Lucene). It communicates with elasticsearch via JSON.

http://pyelasticsearch.readthedocs.org/en/latest/

http://www.elasticsearch.org/

elasticsearch is an Open Source (Apache 2), Distributed, RESTful, Search Engine built on top of Apache Lucene.

The company behind elasticsearch:

http://www.elasticsearch.com/

- Vasudev Ram
www.dancingbison.com

Saturday, November 24, 2012

Mobile Google search with "classic" interface

If you are on a mobile device and want the classic Google search interface instead of the mobile one, use this form:

www.google.com/search?nomo=1&q=your-query-here

-Vasudev Ram
www.dancingbison.com

Thursday, November 15, 2012

Sunday, October 14, 2012

pattern, a Python web mining and NLP tool

By Vasudev Ram


pattern is a web mining and NLP (Natural Language Processing) library for Python.

It is from CLiPS (Computational Linguistics & Psycholinguistics), "a research center associated with the Linguistics department of the faculty of Arts of the University of Antwerp."

From the site:

[ It bundles tools for data retrieval (Google + Twitter + Wikipedia API, web spider, HTML DOM parser), text analysis (rule-based shallow parser, WordNet interface, syntactical + semantical n-gram search algorithm, tf-idf + cosine similarity + LSA metrics), clustering and classification (k-means, KNN, SVM), and data visualization (graph networks). ]

Example usage and output - from the site:
>>> from pattern.web import Twitter, plaintext
>>> for tweet in Twitter().search('"more important than"', cached=False):
>>>    print plaintext(tweet.description)
 
'HINT: The mobile web is more important than mobile apps.'
'Start slowly, direction is more important than speed.'
'Imagination is more important than knowledge. - Albert Einstein'
...
I installed it (download the zip file, extract it and do "python setup.py install"); then tried it out with the above test program and a few variations on it. It partially works; i.e. it's able to fetch some tweets, but in some cases it gives errors that seem to be related to Unicode.

It also has an NLP module for English and a few other languages, plus some other stuff.

UPDATE:

It is now working. Got it to fetch these recent tweets of mine (from my @vasudevram Twitter profile):
IGNORE THIS (testing a Twitter tool). test===444
IGNORE THIS (testing a Twitter tool). test===333
IGNORE THIS (testing a Twitter tool). test===222
IGNORE THIS (testing a Twitter tool). test===111

- Vasudev Ram - Dancing Bison Enterprises

Thursday, August 23, 2012

Fulltext, Python library to convert documents and media to text - for full-text indexing for search


Fulltext is a simple Python library for converting document and media files to text. It's main purpose is for use with full-text indexing systems.

See: https://github.com/btimby/fulltext

and

http://pypi.python.org/pypi/fulltext/0.1-1 (Site giving an error at present)

For example, to easily extract text from a PDF file:

> python
> import fulltext
> fulltext.get('resume.pdf')
'Experience: ...'

Excerpt from the github site for fulltext:

[ Fulltext is a library that makes converting various file formats to plain text simple. Mostly it is a wrapper around shell tools. It will execute the shell program, scrape it's results and then post-process the results to pack as much text into as little space as possible.

Supported formats:
The following formats are supported using the command line apps listed.

application/pdf: pdftotext
application/msword: antiword
application/vnd.openxmlformats-officedocument.wordprocessingml.document:
docx2txt
application/vnd.ms-excel: convertxls2csv
application/rtf: unrtf
application/vnd.oasis.opendocument.text: odt2txt
application/vnd.oasis.opendocument.spreadsheet: odt2txt
application/zip: funzip
application/x-tar, gzip: tar & gunzip
application/x-tar, bzip2: tar & bunzip2
application/rar: unrar
text/html: html2text
text/xml: html2text
image/jpeg: exiftool
video/mpeg: exiftool
audio/mpeg: exiftool
application/octet-stream: strings ]

Inspired by nature.
- dancingbison.com | @vasudevram | jugad2.blogspot.com

Sunday, August 5, 2012

Wow, subtle Google search bug

charles could if he would, james would if he could - Google Search

Try to figure out what the bug is.

(I think it is a bug. Could be wrong.)

- Vasudev Ram
www.dancingbison.com

Wednesday, July 11, 2012

Free Google Online Course - Power Searching with Google


By Vasudev Ram


Seen via an @GoogleAnalytics tweet:

Google is conducting a free online course called "Power Searching with Google".

It will run for a few weeks - the first class is online now - and those completing the two assessments (one mid-term, one final) will get an emailed printable certificate. The course includes the opportunity to hangout (using Google+ Hangouts) with others taking the course and also to ask questions about search to search experts from Google.

Details at the link below. Anyone can sign up for free.

Power Searching with Google

UPDATE: Adding a bit more info about dates and times:

[ Registration for Powering Searching with Google is open until July 16th, however, the first course will be released today (July 10). After that, new classes will become available every Tuesday, Wednesday, Thursday. Attendees will have a two week window to complete them and earn their certificate. ]

So sign up fast if you want to join. There are already a lot of people signed up - I know this because I got some messages like "too many people are accessing this file" for some parts of Class 1.

- Vasudev Ram - Dancing Bison Enterprises

Monday, June 11, 2012

Stand and deliver ...

... er, I meant search and open (*).

Seen on Mike Driscoll's blog via Twitter:

(*) As in, search the web for a phrase and open the result(s) - which is what Mike's post shows how to do, using a few Python libraries - urllibr/urllib2, requests, mechanize and webbrowser. I particularly like the use of webbrowser to automate the opening of the results. The webbrowser module has some options to customize its behavior - see the Python docs for it.

Speaking of webbrowser, spynner looks interesting too.

And the "stand and deliver" bit refers to stories about highwaymen in England of past centuries: http://www.stand-and-deliver.org.uk/ - from memories of historical novels I read as a kid :)

Inspired by nature.
- dancingbison.com | @vasudevram | jugad2.blogspot.com

Wednesday, May 9, 2012

Swiftype trying to build better search for your site

http://m.techcrunch.com/2012/05/08/swiftype-launch/

I find the idea of "crawling your site multiple times, refining the results as they go", interesting.
They should provide some more details though.

Thursday, April 12, 2012

New AWS service - CloudSearch

http://thenextweb.com/insider/2012/04/12/put-some-search-into-your-apps-for-0-12-per-hour-with-new-amazon-cloudsearch-service/

Pay-per-use model, like other AWS services.
Article linked above has more links, including a post by Amazon CTO Werner Vogels about the CloudSearch service.

Friday, March 30, 2012

DuckDuckGo tech goodies

https://duckduckgo.com/tech.html

Jusr seen, one or two seem to have issues, but overall, idea seems good.

- Vasudev

Monday, March 12, 2012

Monday, July 4, 2011

DuckDuckGo search engine may be useful for privacy and security

By Vasudev Ram - dancingbison.com | @vasudevram | jugad2.blogspot.com


DuckDuckGo (*) is a general purpose search engine like Google and Yahoo! search, created by Gabriel Weinberg, who had earlier sold another startup of his (as per a BusinessInsider.com article that I read).

(*) Yes, I know, the name DuckDuckGo is odd. But don't let that deter you from checking it out. The service may have benefits. I also was not too interested in it for a while, though I knew about it from a while ago, and that was partly due to the name.

But recently, I got more interested it it, due to seeing it crawl my web site http://www.dancingbison.com a few times over the last few weeks (as shown by Google Analytics), and then deciding to check it out some more on a whim, and then seeing some interesting stuff about it, namely, that it seems to put more emphasis on privacy (and hence security) than the major search engines like Google. Also, I like the relatively clean, uncluttered user interface of DuckDuckGo.

I also found that there were a handful of positive reviews of DuckDuckGo in well-known publications.

So check it out if you like: http://duckduckgo.com

And maybe, use DuckDuckGo itself to search for more links about it, including the reviews I mentioned above, etc. That may be an interesting test of it, so I leave it as an exercise for the reader (as textbooks are prone to say :-)


- Vasudev Ram - dancingbison.com

Friday, July 1, 2011

Jetslide - geeky queries, mobile and personalized news

By Vasudev Ram - dancingbison.com | @vasudevram | jugad2.blogspot.com

I had tweeted about Jetslide earlier. It's sort of a real-time news site for hacker-related news. Very roughly like hackerstream.com and Peter Cooper (@peterc)'s hackerslide.com, both of which I had also
tweeted about earlier.

Just got an invite to Jetslide.

Excerpts from the email and from the links in it:

[
You can jetslide anonymously but also get some benefits when you login.
...
With Jetsli.de you get the latest news out of twitter, hackernews, delicious, dzone, reddit and digg. The algorithm of jetslide will boost articles with a higher share count (retweets or diggs etc.) and tries really hard to avoid that you'll waste your time e.g. it reduces
duplicates and spam!
...
At the moment the following read modes are available:
Daily: Every 24h Jetslide marks your topics as read.
Auto: Automatically marks your topics as read when you click on the next topic.
...
Read the news on your mobile device like Android.
...
Use geeky queries like e.g. "elasticsearch^2 OR solr" which boosts articles containing elasticsearch.
...
To get posts directly sorted against the share count of one network
you can append the sort parameter, e.g. sort=reddit.com will give you the articles from reddit and sorted against the share count of reddit.
...
Last but not least you can get articles from an url or in this example via domain:
domain:techcrunch.com
]

Seems like an interesting attempt. I think we need more such "geeky query"-type search services :-)

Some time ago I had tweeted (in reply, IIRC, to a tweet by Stormy Peters - @storming), that Gmail should have an SQL-style search facility. (Stop for 1 minute to imagine what you could do with that.) So should other services that provide any kind of search.

The current search facilities, while of some use, are not enough. Of course, initially only technical people would be able to use new facilities such as SQL-style queries, but I don't think it's such a big deal for laymen to learn it. After all, SQL was originally invented - by IBM - with the goal of being an end-user query language (as I had blogged on my earlier blog, jugad's Journal - link below), and though that goal did not really work out - instead, SQL became a programmers' language - nowadays, with vastly greater computer awareness and skills in the general population, I don't see why they should not be able to pick up SQL, at least for basic to intermediate uses.

Earlier blog post about SQL being originally designed for end-users:

IBM Database Connectivity for JavaScript - will the wheel come full circle?
http://jugad.livejournal.com/143125.html

Links about Jetslide and its maker:

http://www.pannous.info/products/jetslide-mobile-and-personal-news4geeks/

http://www.pannous.info/products/

Posted via email.
- Vasudev Ram