Login Register






Document Crawler? filter_list
Author
Message
RE: Document Crawler? #8
(08-28-2014, 02:12 PM)Singularity Wrote: I also want to mention that it would be wise to sort the array by placing the most used words first, and then the least used words. Since he is really wanting to only find the words that are used more often. It's tedious to self scroll through the array list to find the words used most, so making a hierarchy out of it would be helpful.

Something like this in Ruby:
Code:
#!/usr/bin/ruby puts "File: " print '> ' file = File.open(gets.chomp, 'rb') { |f| f.read } puts "Search for words used X or more times" print "X: " input = gets.chomp.to_i word_count = Hash.new words = file.split(' ') words.each do |w| if word_count.include?(w) word_count[w] += 1 else word_count[w] = 1 end end word_count = Hash[word_count.sort_by{ |w, a| a }].to_a.reverse puts "WORD : TIMES USED" word_count.each do |word, amount| if amount >= input puts "#{word} : #{amount}" end end

Allowing the user to only display words used X or more times, and echoing them out in a higher to lower sorted order.

Here is a live example with some tweaked code to make it work on the online interpreter.
(I'm not sure why the sorting is failing on the online interpreter, but I assume it's some sort of issue with the older Ruby version it uses. It probably lacks the sort technique I used)
http://repl.it/Xjc

Edit --
Just seen the graph, so eh this isn't much useful besides the fact of being able to control which words get output by volume.

Would there be a way to get it organized by the number of words, from least to greatest, or vice versa?.

Reply





Messages In This Thread