Tuesday, June 23, 2009

Immigration to foreign countries - some numbers

Leverage the Skills cost arbitrage - Immigrate.
Here are a few notes that I made. Please add to this if you have any other information.

- Improve your SW skills
Browse through the jobs, and look at the common things that are being asked for.
Keep learning the things. Try to apply them in freelance projects.

Applying for US :

find H1B sponsors.
http://corp-corp.com/h1bonlinejobfair_js.htm?gclid=COO7xd2ot5YCFRNPegodi0IZKQ

Foreign VISA consultants :
http://www.visahouse.net/career_guide.php


Online Job Search Sites
Naukri.

http://www.iitjobs.com
http://www.simplyhired.com

Salaries from payscale (software developer):
UAE 1,20,000 AED - 15 lakh INR
Singapore 38000 SGD - 12 lakh INR
Australia 53000 AUD - 17 lakh INR
Canada 75000 CAD - 30 lakh INR
United K 27000 GBP - 22 lakh INR
USA 64000 USD - 31 lakh INR
Switzerland 89000 CHF - 37 lakh INR
Ireland - 34500 EUR - 23 lakh INR
Denmark - 63000 USD - 30 lakh INR

HSMP program needs you to have 2 lakhs in your account

The fees are 350 GBP - 25, 000 INR

--
CANADA:
Canada - Federal Skilled Worker Program
greater than 67 points in the qualifying test - done
At least one year of experience in one of the NOC occupations list:
I come under : 0213 Computer and information system managers :
Software Engineers and Designers (2173)

Immigration Blog :
http://www.canadavisa.com/canada-immigration-blog/

Federal Skilled Worker Program : (New Instructions )
http://www.canadavisa.com/new-instructions-federal-skilled-worker-applications.html

What all are needed for the VISA :
http://www.canadavisa.com/canadian-immigration-faq-skilled-workers.html

4 months prior to the evaluation

Settlement of Funds :
This is the amount that is required to pass through the VISA process (settlement of funds)
10 601 Canadian dollars = 4.2 lakh INR
this is waived if you have arranged employment in Canada
VISA application fees :
22,500 INR (550 CDN)
http://www.canadavisa.com/federal-skilled-workers-processing-fees.html

Arranged Employment :
http://www.canadavisa.com/fast-track-canada-immigration-visa-application.html

The Employer - Employee - Arranged Employment process :
http://www.movetoedmonton.com/foreign/arranged/

VISA processing time :
http://www.canadavisa.com/federal-skilled-worker-processing-times.html#asiaandpacific
Seems to take 72 months. Isnt that too long!
This 72 months is for the Permanent Residentship VISA. You get the Federal Skilled Worker program VISA earlier.

Thursday, June 04, 2009

Pentaho Data Integration - Scalable ETL deployments

Q&A notes from the following webinar for the benefit of PDI users.





Session number:  713773880

Ranadeep Bhattacharya - 11:48 pm

Q: What do you mean by read in parallel? Does that mean only a part of the file is available in each slave?

Matt Casters - 11:49 pm

A: That's exactly what the algorithm does.  It splits the file by size and divides data ranges over the available nodes.

_________________________________________________________________



Ranadeep Bhattacharya - 11:50 pm

Q: But is the file physically located on a single server or split between the 10 or 20?

Matt Casters - 11:50 pm

A: Located on a single shared filesystem.  So the same file is read by N nodes. 

_________________________________________________________________



abhishek manocha - 11:50 pm

Q: So as i understand is their a limitation of clustering only possible if we choose CSV as our input step? I doesnt work on Table Input step?

Matt Casters - 11:52 pm

A: You can do the same thing with a Table Input node, but you need to tweak the SQL statement that is executed since you only want a part of the rows.  Usually it involves using a MOD (%) operator and internal variables representing the node # and # of nodes.

_________________________________________________________________



Robert Folkerts - 11:51 pm

Q: Were there experiments with dimension lookups when populating a fact table?  That is my 'bread and butter' case.

Matt Casters - 11:53 pm

A: Not yet Robert.  With the new cache pre-load option it would make an interesting experiment for sure.

_________________________________________________________________



Dan Jolly - 11:54 pm

Q: When is 3.2 scheduled for GA?

Matt Casters - 11:55 pm

A: Dan, 3.2.0-stable was released last week.

_________________________________________________________________



Vijayaraghavan Amirisetty - 11:56 pm

Q: Does PDI plan to support other partitioning methods like key-range partitioning or hash partitioning in the future ? - Vijay

Matt Casters - 11:58 pm

A: It's not on our roadmap right away.  That being said, it's possible to do now both though Partitioning plugins as well as though a calculation. (simply calculate a partition # and do a mod part on that)

_________________________________________________________________



Peter Schmidt - 11:55 pm

Q: Can you please re-explain the diff between 50/Sort and 100/Sort and 300/Sort.

Matt Casters - 11:56 pm

A: The only difference is the size of the lineitem.tbl file size.  300=1.8B rows, 100=600M rows, etc

_________________________________________________________________



abhishek manocha - 11:57 pm

Q: So considering a scenrio where I have 80 small db inputs and I need to collate them in one central target db, with scehduling of once in a hour (24 times a day) for all sources, clustering make sense?  

Matt Casters - 12:00 am

A: It can make sense if the CPU consumption on your one server is a bottleneck.  If that's not the case, you don't really need to do it.

_________________________________________________________________



sanjeev sagar - 11:59 pm

Q: i joined late but which benchmark it is?

Lance Walter - 12:01 am

A: The whitepaper on bayon-technologies has more details. It uses TPC-H data, but is not a "benchmark" by design.

_________________________________________________________________



Dan Jolly - 12:01 am

Q: Is this cost model based on the EC2 costs?

Lance Walter - 12:01 am

A: yes, computing as well as storage costs on EC2

_________________________________________________________________



prem brahmandam - 11:54 pm

Q: Can we get a sample transform using "table input" step with tweaked query..

Matt Casters - 12:03 am

A: SELECT * FROM foo WHERE mod(id, ${Internal.Step.Unique.Count}) = ${Internal.Step.Unique.Number}

_________________________________________________________________



Peter Schmidt - 12:02 am

Q: If most of your ETL uses table input/table output steps, what changes need to be made to one's transformations, it sounds like if you were reading from flat files, you wouldn't have to do much to configure this to work?

Matt Casters - 12:05 am

A: Peter, it highly depends on the question if your source database can make use of multiple CPUs, etc.  The best strategy is to partition/shard the source and target databases as well. (see also prem's question above)

_________________________________________________________________



Peter Schmidt - 12:08 am

Q: Quick question on the sample transform query, so I am assuming you'd have to put a wrapper around this that increments the count and number)?

Matt Casters - 12:09 am

A: It goes without saying that those internal variables are set automatically in a clustered run.  

_________________________________________________________________



Bret Landon - 12:08 am

Q: Is any of this based on the hadoop methodology?

Matt Casters - 12:10 am

A: No Hadoop cluster is needed although we have plans to make use of Hadoop clusters in the near future.

_________________________________________________________________



Laura Moche - 12:10 am

Q: Were the EC2 servers from this test case dedicated to this test?  Or were the servers shared with other processing?  

Nicholas Goodman - 12:11 am

A: We dedicated the use for the transformations.  But EC2 instances aren't dedicated - they are shared with other EC2 users...

_________________________________________________________________



Dan Jolly - 11:57 pm

Q: What is that top level number

Nicholas Goodman - 12:12 am

A: sorted 450k / rec / sec for 40 nodes

_________________________________________________________________



abhishek manocha - 12:08 am

Q: No Matt, building on the Peter's question, what if the  source database are really on different machines and partitioning/sharding is not an option as in case of mysql 4?

Nicholas Goodman - 12:13 am

A: There are things that can be done to partition the connection on the PDI side.  ie - if you know that host xyz keeps partition 1, and abc keeps partition 2 you can set that up and we'll use just plain 'ole JDBC

_________________________________________________________________



Venu Ambekar - 12:12 am

Q: Is there a capability to handle only incremental changes from a datasource, instead of depending upon the time-stamps of the tables in the datasource.

Nicholas Goodman - 12:14 am

A: Yes.  PDI has capabilities for detecting changes from data sources and you can help you only process those changes.

_________________________________________________________________



Lakshman Bulusu - 12:14 am

Q: What about EL-T in the CLoud?

Matt Casters - 12:16 am

A: It highly depends on the situation.  It still depends on the capabilities of the database(s) (parallelism etc) used.  Suffice it to say that we always recommend you to make that call yourself in PDI.

_________________________________________________________________



abhishek manocha - 12:16 am

Q: Oh I see Partitioning on the PDI side itself, evn if not supported by underlying db ?

Matt Casters - 12:17 am

A: Yes, we refered to data partitioning in the PDI streams earlier, not just database table partitioning.

_________________________________________________________________



Tony Sidhu - 12:13 am

Q: is it practical to run processing in the cloud, that is inputing and outputing to a database that runs in the office?

Matt Casters - 12:21 am

A: Tony, barring any extreme case, I don't think so.  We have another partner that did something similar but keeping the data on the cloud because of cost savings (30%).  It has to be noted that the machine needed to be hosting reports, analyses, etc.  

_________________________________________________________________



Scott Sorensen - 12:21 am

Q: Could you elaborate on 'capabilities to detect changes in a data source' - or provide a reference on how this is done.

Nicholas Goodman - 12:23 am

A: sure... couple of quesitons on this.  There are a few different ways to approach this - none of which is any PDI silver bullet.  There's a step to compare rows (one stream is reference, other is changed) and output diffs.  You can simply parameterize your

Nicholas Goodman - 12:23 am

A: queries so that they only take "update_dt > {last_time_I_checked}"

_________________________________________________________________



Kamal Trivedi - 12:16 am

Q: what if in future the cloud location is moved overseas

Matt Casters - 12:22 am

A: If your infrastructure cloud is moved then it all depends on data volumes, internet speed, etc whether or not you would get in trouble.  With the internet getting faster all the time, I doubt this will become an issue soon.

_________________________________________________________________



Steve McAtee - 12:18 am

Q: Will a recording of this presentation be available offline?



Matthew Papertsian - 12:23 am

A: Yes - the recording will be avilable within 48 hours and sent to you via email



_________________________________________________________________



abhishek manocha - 12:22 am

Q: What's the role of Carte in all this clustering? I was under the impression thats its internal to PDI, if I going to EC2, where does it fit?

Matt Casters - 12:24 am

A: Carte is simply a small webserver that listens to the outstide world.  It can be given transformations, jobs etc to execute. It's controlled remotely.  It's launched during startup of an EC2 host.

_________________________________________________________________



Dan Jolly - 12:20 am

Q: are there any similar case studies?

Nicholas Goodman - 12:24 am

A: Hi Dan - It'd be great if we could get customers to publish some of their own results and case studies for "big data."  I'm not aware of any case studies that look at CLUSTERING/PARTITIONING explicitly.

Lance Walter - 12:24 am

A: here's the best public example - here's the announcement http://www.pentaho.com/news/releases/20090210_nutricia_uses_pentaho_on_amazon_cloud.php  and here is the technical case study http://tinyurl.com/q9b9hs

Nicholas Goodman - 12:25 am

A: PS - How's the weather in Colorado today?

_________________________________________________________________



abhishek manocha - 12:16 am

Q: Can you give more info on that Nick please... on the incremental data?

Nicholas Goodman - 12:26 am

A: See below... there's nothing really "Big Data" special about change data capture.

Nicholas Goodman - 12:27 am

A: there's a variety of techniques.. check the wiki/pentaho training/dev lists for info on how to do this.

_________________________________________________________________



abhishek manocha - 11:58 pm

Q: It may be obvious question, but in last 10 days I have touched the surface of PDI only and done sample test on single workstation of mine

Matthew Papertsian - 12:27 am

A: Abhishek - can you please complete your question as I am not certain what you are asking?

_________________________________________________________________



sanjeev sagar - 11:59 pm

Q: or which tools were used for these fig.?

Nicholas Goodman - 12:28 am

A: This was PDI 3.2 (pre GA it was inbetween release candidates).  The exact build # is in the whitepaper.

_________________________________________________________________



abhishek manocha - 12:28 am

Q: hi Matt, so you dont really recoment it for office use as you mentioned to tony?

Matt Casters - 12:30 am

A: Well, given the fact that you now have "local" clouds in large corporation and that virtualization keeps growing, I might be completely wrong.  Then again, it was a very specific question that Tony had.

_________________________________________________________________



abhishek manocha - 12:30 am

Q: Ok, fair enough Matt

Matt Casters - 12:31 am

A: Sure thing!

_________________________________________________________________



Ulrich Riedel - 12:30 am

Q: I have seen in the webcast that sorting 1 billion lines a month costs app. 32,000$. Why is this price said to be cheap? Are there comparable prices?

Matt Casters - 12:32 am

A: It's only 4$ to sort a billion rows.  I think the situation was that if you needed to do it every hour or so, you would spend that money.  Cloud works best economically if it takes care of peak loads.

Matt Casters - 12:32 am

A: Sorry 6$ :-)

Thursday, June 05, 2008

Payback (in km) when you go for a second hand bike

I decided to go for a second hand bike as I do not prefer to go on a bike in city traffic.
I intend to use the bike for long distance drives on weekends. As you see in the calculations below,
one recovers the cost of a new pulsar in 24,500 km while a second hand pulsar's price is recovered in 16,000 km. The maintainance cost has been factored in the calculations ( it is twice the amount that a regular bike would have).

btw, I bought this bike from bike galaxy at basappa circle on lalbagh road. There are a lot of second hand bike dealers (agents in fact) near minerva circle and VV puram. The same bike would have cost be 3 to 4 thousand cheaper if I would have bought it from the seller directly. I had to pay 1000 Rs to the agent to buy this bike.
The agent gathers this bike from exchange melas, direct sellers, etc.
How far my calculations are correct?
Only time will tell :)

Monday, June 02, 2008

Bought a bike. Ready to go

The bike, a 2003 pulsar (non DTSI , non alloy wheels, non kickstart) cost be 28,600 (This includes 750 Rs as insurance, 1000 Rs as commission, 450 Rs as registration charges). I bargained with the owner o reduce the price from 29,000 to 26,600.
The odometer reading has been manipulated, so I don’t know how many KM it has run. No idea of the mileage (trying to figure that out now), the RPM meter malfunctions and the idlng RPM has been set to 4,000 (2,200 should be ideal).
Meter has started, and I am in the process of learning the bike internals at present. I intend to use this mike for long distance rides, and the target would be to make a trip to Kolhapur from Bangalore (600 km one way on NH4).

Actually, it is always better to buy a bike directly from the owner as you can save the commission. But you can do this only if you have the ability to gauge the condition of the bike and judge the price of the bike correctly. I feel that I paid a bit more for the bike (I had started with a budge of 25,000!), but since I have already bought it, it is okay.

Here are a few tips that I found useful:

Mouthshut

60kph

btw, I bought the bike from Minerva circle. You will find a lot of shops selling used bikes there. Do check for the Road tax papers and the RC book before buying the bike. It is always good to take along a friend who is knowledgeable about bikes.


Sunday, June 01, 2008

Yahoo Ad "Sense" screwed up


I come to office on a monday morning, log on to Yahoo Mail and notice something
different. The banner ad was in Chinese! (Could be Japanese too). I am trying to understand the logic behind being presented this ad! There are no Chinese e-mails in my mailbox, no Chinese girlfriends. The only thing that is vaguely related to Chinese lying in my inbox is the eBay correspondence about the purchase of a Mobile Phone (CECT make) which is manufactured in china.
Funny!

PS: I use google Adsense on my blog, and most of the times, the textual ad's that are displayed related to the context of the blogpost.

Sunday, December 16, 2007

Image Matching

I had been toying around with the idea on extending Video search to include the audio and video (image) data associated with the video in addition to the user specified tags. This is one of my weekend projects, others being the semantic representation of knowledge and android app development.
This weekend I spent some of my time get out a rough cut of the "Image Indexing using Color Correlograms" [the same with my notes] which could be helpful in my Video Search project. The principle behind an image correlogram [Correloation + histogram] is the spatial correlation between the image colors. A correlogram can be defined as f(c1,c2,k), where f is the number of pixels of color c2 around a pixel of color c1 which are at a distance of k from the c1 colored pixel.
The results looked promising. At present I used only the G channel to generate the correlogram, and I would be refining it over time.
Some results : I chose come images obtained by the query: road via google image search, calculate the correlograms for each and then used the metric described in the paper to find the closeness. The number adjacent to each image describes how close it is to Image 1. Lower the number the closer the image is to Image1

Image 1:





Image 2: 16.32
Image 3: 4.84 Image 4: 9.51

Image 5: 15.06 Image 6: 17.03Image 7: 15.05

The closest images to Image 1 are Image 3 and Image 4, which looks intuitive.
Image 5 and Image 7 have similar correlograms which also seems fine.
But the observation that Images 2,5,6 and 7 are at almost equidistant to Image 1 is not very palpable. Especially as Image 2 is no way related to Image 1. On the other hand, it is not possible to draw inferences on the robustness of this approach with such a small set of test images. I'd be do some more research and come up with enhancements to get better results.

At present I just used the green channel of the image in the creation of the correlogram as histogram of the green channel is closet to the luminosity histogram. [ The reason for this is that the human eye is more sensitive to the color green than any other colors]. The current algorithm calculates f using all the pixels in an image and computing f is quite expensive O(n^2*d) . [The complexity has been mentioned to be O(n^2*d^2) in the paper and I will try to clarify this with the authors]. The space complexity is O(m^2 * d) [where m is the number of possible colors] Further improvements in time and space can be made by calculating the correlogram for select regions of the image (high gradient blocks).

btw, it took me quite some time to set up a development environment in windows.
I used the opencv sdk for reading the images, eclipse IDE with CDT, and mingw (gcc for windows)for the compiler.


Some Informative links that I followed:
opencv
http://www.site.uottawa.ca/~laganier/tutorial/opencv+directshow/cvision.htm
http://www.cs.iit.edu/~agam/cs512/lect-notes/opencv-intro/opencv-intro.html#SECTION00041000000000000000
http://www.xpercept.com/opencv.htm

eclipse cdt:
http://www.cs.umanitoba.ca/~eclipse/7-EclipseCDT.pdf

updates:
Don Dodge on video search

Saturday, November 17, 2007

Autorickshaw economics

Today, while I was returning from a Shopping mall in bangalore in an Autorickshaw {Autorickshaws are 2 stroke, 3 wheelers which are a common mode of transportation in India. [Equivalant of taxi's in the US]} the autorickshaw driver said something to me in Kannada [local language]. I couldnt understand it, but, out of curiosity, I said "Nannage Kannada swapla swalpa gotthu; Hindi, English, Telugu". {this is what I usually say when someone talks to me in Kannada - I know very little Kannada. These are the languages I know}
The driver (about 25-30 years in age) then told me that he had been waiting for a passenger from 1 o'clock, and he did not get a single passenger till the time I asked him to take me home [It was 3:15. He had been idle for 2.25 hours!]. He also tole me that, he would have gone home if I hadn't boarded his auto.
The driver seemed friendly, hence I asked him some more questions. The information that I got is as follows:
mileage of an auto - 20 km per litre of petrol (gas). This costs him 2.5 Rs per km.
The standard charge is 6 Rs per km, thus he makes a profit of 3.5 Rs per km.
He travels close to 80 km every day [excluding the "idle" rides]. Thus, he earns around 280 Rs [7 USD] per day i.e. 2,600 USD per annum . That is pretty low as compared to and average engineer's salary [12000 USD per annum].
I can help these guys by getting them passengers. Using mobile phones to connect the prospective passengers to the idle auto rickshaws. I'll post the details of my idea soon.
Now, back to android!

Tuesday, October 23, 2007

I like this, I'd find it there

WHAT DO I LIKE:
1. I'd like to do something where I could always learn something new.
- Travel [learn history / geography / cultures]
- Photography
- Reading Papers / implementing them
- Reading Blogs / Google Videos

2. Something that I feel that I can make a difference
- Semantic web / agents
- 3D Reconstruction
- President [MEA events / publicity]
- SEEK Teaching

3. Something where I could motivate others / teach others
- SEEK
- TA ship
4. Something where I find solutions to real problems
- planning a tour
- building an mp3 jack to plug in the iPod to the car system
- finding out a way to have the minimum trips to dispose garbage

All this with enough money so that I could pursue my interests and hobbies [travel / photography / etc ]

Where can I find all these ?
Working on Innovative / hard (not necessarily) and open ended projects
- Research Labs in Academic Institutes
- Open Source Projects
- (Google) Research Labs [Rich]

What do I do to get what I want.

Sunday, August 05, 2007

Personal Productivity

I thought that it was time to rethink over the way I work, prioritise tasks and implement them.
Hence, I read up a few things on the web.
Found a couple of good resources and suggestions.

I have bookmarked the links at:

http://del.icio.us/amirivija/Productivity

I also subscribed to a couple of blogs on productivity.

Here are a few notes I made for myself based upon what all I read:

Things to implement:
1] Keep a record of the start and end times of different activities
2] Cut down e-mail alerts. cut down IM. Look at the e-mails only 4 times a day. [Morning, After Lunch, Before Leaving - set aside time for rss feeds ets]
3] Clarify objectives, before starting any work
4] Break down the big problem into smaller ones

5] Work tends to expand in the time available.. [so compartmentalize the work] Hence, shrink the time you are going to work.. [ Plan to work only for 5 hours per day]

6] Stop multi tasking..

7] Do the task that gives the most benefit.

8] Join a group to keep you motivated
1] Apping group
2] Build an online programming network

9] Picture of a goal : expert computer science engineer

10] be patient

11] seek inspiration - blogs, people , articles

12] never skip anything two days in a row

13] Apply the ""do it, delegate it, defer it, drop it"" rule for all the stuff that' clogging your system

I also stumbled across the book "Getting Things done" by david allen. I would like to read this book too.

This one also has pointers to some good stuff
http://zenhabits.net/2007/02/beginners-guide-to-gtd/

And last but not the least, I reorganised my Google reader feeds

Saturday, August 04, 2007

I am a Kinesthetic Learner

This is what http://www.metamath.com/multiple/multiple_choice_questions.html has to say about my learning style.

The results of Amirisetty Vijayaraghavan's learning inventory are:

Visual/Nonverbal 30 Visual/Verbal 28 Auditory 24 Kinesthetic 36

Your primary learning style is:

The Tactile/ Kinesthetic Learning Style


You learn best when physically engaged in a "hands on" activity. In the classroom, you benefit from a lab setting where you can manipulate materials to learn new information. You learn best when you can be physically active in the learning environment. You benefit from instructors who encourage in-class demonstrations, "hands on" student learning experiences, and field work outside the classroom.

Strategies for the Tactile/ Kinesthetic Learner:

To help you stay focused on class lecture, sit near the front of the room and take notes throughout the class period. Don't worry about correct spelling or writing in complete sentences. Jot down key words and draw pictures or make charts to help you remember the information you are hearing.

When studying, walk back and forth with textbook, notes, or flashcards in hand and read the information out loud.

Think of ways to make your learning tangible, i.e. something you can put your hands on. For example, make a model that illustrates a key concept. Spend extra time in a lab setting to learn an important procedure. Spend time in the field (e.g. a museum, historical site, or job site) to gain first-hand experience of your subject matter.

To learn a sequence of steps, make 3'x 5' flashcards for each step. Arrange the cards on a table top to represent the correct sequence. Put words, symbols, or pictures on your flashcards -- anything that helps you remember the information. Use highlighter pens in contrasting colors to emphasize important points. Limit the amount of information per card to aid recall. Practice putting the cards in order until the sequence becomes automatic.

When reviewing new information, copy key points onto a chalkboard, easel board, or other large writing surface.

Make use of the computer to reinforce learning through the sense of touch. Using word processing software, copy essential information from your notes and textbook. Use graphics, tables, and spreadsheets to further organize material that must be learned.

Listen to audio tapes on a Walkman tape player while exercising. Make your own tapes containing important course information.

Thursday, July 26, 2007

to READ

http://code.google.com/edu/content/submissions/uwspr2007_clustercourse/listing.html

Monday, July 09, 2007

Another Idea - Speech recognition

2-3 years ago I had looked up some work on speech to text conversion.
The approaches were based on Neural networks, and I did not find them helpful.

I dont know the as it is now, but here is the Idea that sprung to my mind yesterday.

Saturday, July 07, 2007

Towards Semantic web. One step at a time

I was traveling back from my hometown to bangalore when I came across this good science fiction article which revolves around semantic web and embedded devices.
Just reinforcing my vision of a world with information at the time we want and at the place we need. Basically it is the availability of information when needed.

I had been reading a few papers on Information retrieval, the semantic web [collaborative filtering/ social networks]. Few of these papers have fascinated me and I believe that they are going to make the world a easier place to live in.

Here is my iota of contribution towards this end:
This is a continuation of my efforts started here.

I implemented the concepts mentioned in the paper above to come up with a program capable to developing a concept map for a small document corpus.

Here is the concept map:


I used the Stanford NLP parser to extract the nounphrases from the document.
JGraphT to represent the graph (I used the DirectedGraph) and JGraph to display the graph.

Lots of improvements to be done:
  1. At present all the document processing and generation of the concept map is done online. That sucks up a lot of memory. Hence I'll have to use persistent storage store the nounPhrase - Document association, the Document information, the nounPhrase - nounPhrase association. I am planning to try hsqldb for this. It supports inmemory databases.
  2. Need to speed up the execution by having some parallel processing. I will try getting help from a professor at IISc [SERC] for this.
  3. I need to add the relationship between two entities (alongwith their weightage). For example: Google has a relationship with yahoo. Right now my program just tells that there is an association. but what type of association [I can be competitors/ successful startups/ young CEO's/ web companies/ great places to work at/etc etc]
  4. Optimise the code.
And I want to learn developing applications for embedded devices too.
I will start with image processing in mobile phones. I bought a second hand nokia 6600 for this.
The two thing I have in mind for mobile applications are:
  • An image processing program that will tell me the destination of the bus when I click the photograph of the destination board of the bus (which is in Kannada) so that I dont have to rush to the bus to ask "Bhaiya, majestic jayega?"
  • An image processing program (again) which converts my notes (which are quite random) into a good power point presentation (It should at least capture the shapes and the content)
Hmm.. seems that I talk a lot. Let me get back to work!

Saturday, June 16, 2007

Distance Learning MS in CS

This one looks good:

http://www.grad.iit.edu/bulletin/programs/cs.html#mast-sci-comp-sci

Cost: 20 760 U.S. dollars = 8,47 ,796.79 Indian rupees


Programming core courses
CS 522 Data Mining
CS 525 Advanced Database Organization
CS 529 Information Retrieval
CS 540 Syntactic Analysis of Programming Languages
CS 546 Parallel Processing
CS 551 Operating System Design and Implementation

Systems core courses
CS 542 Computer Networks I: Fundamentals
CS 544 Computer Networks II: Network Services
CS 547 Wireless Networks
CS 550 Advanced Operating Systems
CS 555 Analytic Models and Simulation of Computer Systems
CS 570 Advanced Computer Architecture
CS 586 Software Systems Architectures

Theory core courses
CS 530 Theory of Computation
CS 532 Formal Languages
CS 533 Computational Geometry
CS 535 Design and Analysis of Algorithms
CS 536 Science of Programming
CS 538 Combinatorial Optimization

Sunday, June 03, 2007

The germ

The 16th International World Wide Web conference was held at Banff, Alberta Canada from May 8 to 12th this year.
The web has revolutionalized the way information flows, the way people learn, the way people do their everyday things, the way people socialize and a zillion other things. In short, it has revolutionized the way we live.
And each year the www conference sets the stage to accelerate this revolution in the years to come.

Here are a few papers that I'd like to read up in the limited time I have.
I collected these docs by searching for pdf site:http://www2007.org/ with google.
I have skimmed through the abstracts of the interesting papers of the first 90 results
I found around 29 papers interesting.

Track: Semantic Web (2)

1] Analysis of Topological Characteristics of Huge Online *

Social Networking Services

http://www2007.org/papers/paper676.pdf

2] Combating Spam in Tagging Systems [Stanford]

http://www2007.org/workshops/paper_97.pdf


Track: Search, Information organization, retrieval and processing (17)

1] Navigation-Aided Retrieval

http://www2007.org/papers/paper162.pdf

2] Supervised Rank Aggregation *

http://www2007.org/papers/paper286.pdf

3] Functional Faceted Web Query Analysis [NSU, Singapore] *

http://www2007.org/workshops/paper_44.pdf

4] Efficient Search in Large Textual Collections *

with Redundancy

http://www2007.org/papers/paper800.pdf

5] Formalization, User Strategy and Interaction Design:

Users’ Behaviour with Discourse Tagging Semantics

http://www2007.org/workshops/paper_30.pdf

6] Sponsored Search with Contexts [upenn]

http://www2007.org/workshops/paper_83.pdf

7] Towards a Semantic Knowledge Base for Yeast Biologists

http://www2007.org/workshops/paper_130.pdf

8] Detecting Near-Duplicates for Web Crawling [google] *

http://www2007.org/papers/paper215.pdf

9] Robust Methodologies for Modeling Web Click [yahoo]*

Distributions

http://www2007.org/papers/paper056.pdf

10] Random Web Crawls[criteo]

http://www2007.org/papers/paper339.pdf

11] An Adaptive Crawler for Locating Hidden-Web Entry Points [Utah]*

http://www2007.org/papers/paper429.pdf

12] Web Projections: Learning from Contextual Subgraphs of the Web

[Microsoft, cmu]*

http://www2007.org/papers/paper551.pdf

13] The Discoverability of the Web [yahoo]*

http://www2007.org/papers/paper592.pdf

14] Extraction and Search of Chemical Formulae in Text

Documents on the Web [PSU]

http://www2007.org/papers/paper100.pdf

15] Summarizing Email Conversations with Clue Words [ucb, cancada] *

http://www2007.org/papers/paper631.pdf

16] Answering Relationship Queries on the Web [IBM]*

http://www2007.org/papers/paper058.pdf

17] Do Not Crawl in the DUST: Different URLs with Similar Text

http://www2007.org/papers/paper194.pdf

Recommender systems, Collaborative filtering (4)

1] Improving Ontology Recommendation and Reuse in

WebCORE by Collaborative Assessments

http://www2007.org/workshops/paper_8.pdf

2] The Complex Dynamics of Collaborative Tagging [Princeton] *

http://www2007.org/papers/paper635.pdf

3] Scaling Up All Pairs Similarity Search [google] **

http://www2007.org/papers/paper342.pdf

4] Applying Collaborative Tagging to E-Learning

http://www2007.org/workshops/paper_56.pdf

Data Mining: (2)

1] Towards Domain-Independent Information Extraction

from Web Tables

http://www2007.org/papers/paper790.pdf

2] Page-level Template Detection via Isotonic Smoothing [yahoo]* omi

http://www2007.org/papers/paper588.pdf

Miscellaneous but interesting (4)

1] Optimal Audio-Visual Representations

for Illiterate Users of Computers

http://www2007.org/papers/paper764.pdf

2] A Mobile Application Framework for the Geospatial Web

http://www2007.org/papers/paper287.pdf

3] Communication as Information-Seeking: The Case for

Mobile Social Software for Developing Regions [Washington]

http://www2007.org/papers/paper669.pdf

4] Tag-Cloud Drawing: Algorithms for Cloud Visualization

http://www2007.org/workshops/paper_12.pdf

Friday, May 25, 2007

Know thy self

I came across this blog post and found it very insightful.

I could relate myself to the author of this post.

A few points that I found interesting. [verbatim from the post]:

-
Beware the traps of justification and attribution when thinking about your life. Many people justify past decisions because they want them to make sense and feel better.

^-- This is what I have been doing the past one year! Fooling myself.

Yes, most clouds have silver linings. But sometimes getting hit in the head with a shovel is just getting hit in the head with a shovel. It’s not fun, there’s not much educational value, and your life really is better without it.

-
I really don't know what I want.
The author gives simple and effective tips to understand what you want:
1. A feel good list <-- In this, you note down the events that made you feel good. eg: understanding the paper on semantic search. eg: Explaining Sudhir on why I feel that semantic search would be better than textual search. eg: Telling a colleague that google patent search would be better than delphion.
eg: Helping a colleague finding a solution to the problem he was facing.


2. An ideas list <-- In this you note down the ideas that come to your mind. The author mentioned sometimes he gets a great idea in the sleep and he just gets up and notes that idea down. This happened to me a couple of times :). 3. Finally, get external inputs. Understand your personality and get inspired. I took up this personality test on: and found that : I was and ENJF http://typelogic.com/enfj.html

ExtravertedIntuitiveFeelingJudging
Strength of the preferences %
33381211

I'd like to read it again.
In short, it says that people like me like to find solutions, look for greener pastures, help others and are pedagogues of humanity.
That is true I believe, but it is for the people around me to tell me that.

Here is something from a test that I took up previously.


More good talk. From the same blog:


This talk is good:

Richard St. John: Secrets of success in 8 words, 3 minutes

  • Do what you love: Passion
  • Determination: Hard work, focus, push yourself, get good in one thing
  • Make something useful: Ideas, serve others
  • Keep your chin up: Persist

I need to look at the highlighted points

One needs to persist in the face of failure and other CRAP:

  • Criticism
  • Rejection
  • Jerks (or another word for a non-nice person)
  • Pressure

Tuesday, May 15, 2007

Personality development for Kids

Here is something I wrote on a piece of paper.. This was an effort for the kids of SEEK, VISHWAS, Bangalore. A part of it has been implemented





Wednesday, May 02, 2007

Syntactic to Semantic Web. The search engine that "Understands" what you want!

I have been looking for current work on text mining, knowledge representation/extraction from unstructured data. I have bookmarked

I found this paper particularly interesting:
"Knowledge Discovery from Semi-Structured Data for Conceptual Organization"


The authors talk about
  • creating a concept map (a graph of co-occurring [in other terms, related] concepts [noun phrases]) from a corpus,
  • then extracting cliques of a particular concept [using a topological sort],
  • and then assigning the documents to that particular clique

What you get a the end is the mapping of a document to concepts. In short, what are the noun phrases that describe the document.
These noun phrases may not always contain all the noun phrases that occur in the document [eg. red soil, which may have occurred in one of the documents containing the concept "rose", but it did not co-occur in any other document containing "rose"]
On the other hand, these noun phrases may include some concepts which were not obvious from that particular document [eg. a document containing "school kids" may not contain the concept "chocolate", but these two concepts have co-occurred in a significant number of documents in the corpus]..

I liked the paper, the concept is very similar to what we do in real life.
We have concepts stored inside our brain [connection between neurons?]. To extract the concept map stored in your brain, you just need to think about all the things that come to your mind when you think about a concept say "s e x".
When ever we come across something new (a new concept), we just associate it to an already existing concept map in our brain. These concepts are in turn linked to information about them (related documents?)

We can even take this a step ahead by assigning weights to the edges [based upon how frequently they co-occur, label the edges[and nodes] with all possible verbs that connect them ]

btw, the authors the Stanford Parser to extract the noun phrases.

you can try it online here.
Moving from syntactic to semantic.. arent we?

Saturday, April 21, 2007

why are we [Indians] prosperous but not happy

This post gives a good idea..
I'd like to read it once more in my leisure and ponder over it.

http://harmanjit.blogspot.com/2007/04/notice-period-to-god.html

Thursday, April 19, 2007

Portal for Kids

I feel that this idea would be a hit..

resources here
http://www.memory-key.com/Parents/strategies.pdf
http://www.journal.naeyc.org/btj/200407/OnlineAndPrintArtResources.pdf
http://www.surreymuseums.org.uk/somethingelse/psycho.htm