Wednesday, December 12, 2018

Automatic writing with Deep Learning: Progress

This is a continuation of the post https://dmitrykan.blogspot.com/2018/05/automatic-writing-with-deep-learning.html. This item was reblogged at Writer's DZone: https://dzone.com/articles/automatic-writing-with-deep-learning-progress

Fast forward few months (apologies for the delay) I can share some findings.
Again, I think, we should take AI co-writer exercises with a grain of salt. However, during this time I have come across practical usage example areas for such systems.

One of them is augmentation of a news article writer. More specifically, when writing a news item, one of the most challenging tasks is to coin a catchy title. Does the title have some trendy phrases in it? Or perhaps it mentions an emerging topic, that captures attention at this given moment? Or reuses a pattern that worked well for this given author? Or just spurs an idea in the author's head?

Copyright: https://www.rogerwilco.co.za/blog/robot-writers-how-ai-will-affect-copywriting


In the following exercise I have set a very modest goal: train a co-writer on previously written texts with an attempt to suggest something useful from them. I could imagine, that this could be extended to texts that are trending or a collection of particularly interesting titles. What have you.

To train such a model I have used Robin Sloan's RNN writer: https://github.com/robinsloan/rnn-writer. The goodies of the project are:
  • Trained on Torch. Nowadays, Torch is leveraged via PyTorch, a deep learning Python library that is nearing its production readiness time.
  • The trained model gets exposed into an Atom -- pluginable editor (I'd imagine, real writers would want to have the model integrated into their favourite editor, like Word).
  • API is available too to integrate into custom apps (and this is exactly how it is integrated with Atom).

I will skip the installation of Torch and training the network and proceed to examples. The rnn-writer github repository has a good set of instructions to proceed with. I have installed Torch and trained the model on a Mac.

First things first: RNN trained on my Master's Thesis "Design and Implementation of Peer-to-Peer Network" (University of Kuopio, 2007).


The text of the Master's Thesis is about 50 pages in English with diagrams and formulas. On one hand, having more data makes NNs learn more word representations and should have larger probability space to predict next word given the condition of the current word or phrase. On the other hand, limiting the input corpus to phrases that have certain domain goal, like writing an email, could leverage a clean set of phrases that a user employs in many typical email passages.

As I got an access to Fox articles, I thought, this could warrant another RNN model and a test. Something to share next time.

Sunday, May 6, 2018

Automatic writing with Deep Learning: Preface


This article was also reblogged at: https://dzone.com/articles/automatic-writing-with-deep-learning-preface


Quite many machine and deep learning problems are directed at building a mapping function of roughly the following form:


Input X ---> Output Y,


where:

X is some sort of an object: an email text, an image, a document; 

Y is either a single class label from a finite set of labels, like spam / no spam, detected object or a cluster name for this document or some number, like salary in the next month or stock price.

While such tasks can be daunting to solve (like sentiment analysis or predicting stock prices in realtime) they require rather clear steps to achieve good levels of mapping accuracy. Again, I'm not discussing situations with lack of training data to cover the modelled phenomenon or poor feature selection.

In contrast, somewhat less straightforward areas of AI are the tasks that present you with a challenge of predicting as fuzzy structures as words, sentences or complete texts. What are the examples? Machine translation for one, natural language generation for another. One may argue, that transcribing audio to text is also such type of mapping, but I'd argue it is not. Audio is a "wave" and the speech detection is an okay solved task (with state of the art above 90% of accuracy), however such an algorithm does not capture the meaning of the produced text,  except for where it is necessary to do the disambiguation of what was said. Again, I have to make it clear, that audio->text problem is not at all easy with its own intricacies, like handling speaker self corrections, noise and so on.



Lately, the task of writing texts with a machine (e.g. here) caught my eye on twitter. Previously, papers from Google on writing poetry or other text producing software were giving me creepy feelings. I somehow undermined the role of such algorithms in the space of natural language processing and language understanding and saw only diminishing value of such systems to users. Again, any challenging tasks might be solved and even bring value to solving other challenging tasks. But who would use an automatic poetry writing system? Why would somebody, I thought, use these systems -- just for fun? My practical mind battled against such "fun" algorithms. Again, making an AI/NLProc system capable of producing anything sensible is hard. Take the task of sentiment analysis, where it is quite unclear what the agreement between experts is, not to mention non-experts.

I think this post has poured enough of text onto the heads of my readers. I will use this post as a self-motivating mechanism to continue the research with systems producing text. My target is to complete the neural network training on the text from my Master thesis and show you some examples for your judgement of the usefulness of such systems.

Saturday, May 5, 2018

AI for lip reading

It is exciting to push your imagination for where else can you apply AI, machine learning and most certainly -- deep learning, that is so popular these days. I came across this question on quora that provoked me to think a bit how would one go about training a neural network to lip read. I don't actually know what made me answer this question more: that found myself in an unusual context sitting on an Angularjs meetup at Google offices in New York City (after work, usual level tired) or the question itself. Whatever the reason, here is my answer:

Source: http://theconversation.com/our-lip-reading-technology-promises-to-make-hearing-aids-more-human-45166

I would probably first start with formalizing what is lip reading process from a human understandable algorithm point of view. May be it is worth to talk to a professional, like a spy or something. Obviously you need training data. Understanding, what is lip reading from the algorithm perspective will affect on what data you need.


    1. To read a word of several syllables you’d need a sequence of anchor lip positions, that represent syllables. Or probably vowels / consonants. See, I don’t know, which one is best. But you’d need to start with the lowest level possible out of which you can compose larger sequences, like letters -> syllables -> words. Let’s call these states.
    2. A particular lip posture (is that the right word?) will most probably map to ambiguous states.
    3. Now the interesting part is how to resolve the ambiguities. Number 2 produces several options. Out of these you can produce a multitude of words that we can call candidates.
    4. Then you need to score candidates based on some local context information. Here it turns into a natural language understanding.
    5. I'd start with seq2seq.

    Tuesday, January 16, 2018

    New Luke on JavaFX

    Hello and Happy New Year to my readers!

    I'm happy to announce release of completely reimplemented Luke -- using JavaFX technology.  Luke is the toolbox for analyzing and maintaining your Lucene / Solr / Elasticsearch index on low level. 

    The implementation was contributed by Tomoko Uchida, who also did the honors of releasing it.

    The excitement of this release is supported by the fact, that in this version Luke becomes fully compliant with ALv2 license! And it gets very close to be contributed to Lucene project. At this point we need lots of testing to make sure JavaFX version is on par with the original thinlet based one.

    Here is how load index screen looks like in new JavaFX luke:


    After navigating to the Solr 7.1 index and pressing OK, here is what luke shows:


    I have loaded an index of Finnish wikipedia with 1,069,778 documents, and luke tells me that the index does not have deletions and was not optimized. Let's go ahead and optimize it:




    Notice, that on this dialogue you can request only expunging of deleted docs, without merging (the costly part for large indices). After optimization's complete, you'll have a full log of actions in front of you to confirm the operation was successful:


    You could also opt for checking the health of your index via Tools -> Check index menu item:



    Let's move to the Search tab. It has changed slightly in that search box has moved to the right, while search settings and other knobs were moved to the left.

    Thinlet version:


    JavaFX version:



    It is more intuitive UI now in terms of access to various tools like Analyzer, Similarity (now with access to parameters of new BM25 ranking model, that became default in Lucene and default in luke) and More Like This. There is a new Sort sub-tab that lets you choose a primary and secondary field to sort on. Collectors tab however is gone: please let us know, if you used it for some task -- would love to learn.

    Moving on to the Analysis tab, I'd like to draw your attention towards really cool functionality of loading custom jars with your implementation of a character filter, tokenizer or token filter to form your custom analyzer. Test these right in the luke UI without the need to reload shards in your Solr / Elasticsearch installation:



    Last, but not least is Logs tab. Essentially you should have been missing it for as long as luke exists: getting a handle of what's happening behind the scenes during an error case or a normal operation.

    In addition, this version of Luke supports the recently released Lucene 7.2.0.

    Wednesday, November 1, 2017

    Will deep learning make other machine learning algorithms obsolete?

    The fourth (fifth?) quoranswer is here! This time we'll talk a bit about deep learning and its role in making other state of the art machine learning methods obsolete.


    Will deep learning make other machine learning algorithms obsolete?


    I will try to take a look at the question from the natural language processing perspective.

    There is a class of problems in NLProc, that might not be benefited from deep learning (DL), at least directly. For the same reasons, machine learning  (ML) cannot help so easily. I will give three examples, which share more or less the same property so hard to model with ML or DL:

    1. Identifying and analyzing a sentiment polarity oriented towards a particular object: person, brand etc. Example: I like phoneX, but dislike phoneY. If you monitor the sentiment situation for the phoneX you'll expect this message to be positive, while negative polarity for the phoneY. One can argue, it is easy / doable with ML / DL, but I doubt you can stay solely within that framework. Most probably you'll need a hybrid with rule-based system, syntactic parsing etc, which somewhat defeats the purpose of DL: be able to train neural network on a large amount of data without domain (linguist) knowledge.

    2. Anaphora resolution. There are systems that use ML (and hence DL can be tried?), like BART coreference system , but most of the research I have seen so far is based around some sort of rules / syntactic parsing (this presentation is quite useful: Anaphora resolution). There is a vast application area for AR, including sentiment analysis and machine translation (also fact extraction, question-answering etc).

    3. Machine translation. Disambiguation, anaphora, object relations, syntax, semantics and more in a single soup. Surely, you can try to model all of these with ML, but commercial systems in MT are more or less done with rules (+ml recently). I'm expecting DL can produce advancements in MT. I'll cite one paper here that uses DL and improves on phrase-based SMT: [1409.3215] Sequence to Sequence Learning with Neural Networks Update: some recent fun experiment with DL based machine translation.

    The list can be extended to knowledge bases etc, but I hope I made my point.

    Sunday, October 29, 2017

    More fun with Google machine translation

    Having posted in quoranswer tag specifically on machine translation tricks and challenges + looking at some fun with Mongolian->Russian translation with Google, I decided to experiment with Mongolian->English pair. To make this work, you'd need a Cyrillic keyboard and type only Russian letters 'а' as input on Mongolian language side. Throughout the text I'll refer to Google Translate as "neural network" or "network", as it has been known that Google has switched its translation system over to a Neural Network implementation.

    So let's get going. It all starts rather sane:



    а   -> a
    аа -> ah

    And as we stack up more letters on the left, we start getting more interesting translations:

    ааа -> Well
    аааа -> ahaha
    ааааа -> sya
    аааааа -> Well
    ааааааа -> uh

    and skipping a bit:

    ааааааааа -> that's all

    (at this point you'd imagine that deep neural network had some fun you teasing it and wants you to stop. But no).

    аааааааааа -> that's ok
    аааааааааааааа -> that's fine

    ааааааааааааааааа -> everything is fine

    ааааааааааааааааааа -> it's a good thing


    And a bit more letters stacked up, the network begs to stop again, threatening:

    ааааааааааааааааааааааааааааааааааааа -> it's all over

    Then, after having enough of statements, the network starts asking questions.

    ааааааааааааааааааааааааааааааааааааааааа -> is it a good thing?

    and answers own question:

    аааааааааааааааааааааааааааааааааааааааааа -> it's a good thing

    few comments here and there:

    ааааааааааааааааааааааааааааааааааааааааааааааааааа -> a good time

    аааааааааааааааааааааааааааааааааааааааааааааааааааа-> to have a good time

    Eventually, more dictionary entries crop in:

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> to a whirlwind

    And, unexpectedly:

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> to make a date
    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> to make a living

    Then, the network starts to output:

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> to make a dicision

    And begs me to put some sane words in instead of the letter non-sense:

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> put your own word

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a whistle-blower

    The latter one is probably meant as an offence to add colour to network's ask.

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> have a private time in the world

    Notice how general words are, like "private", "time", "world". Still they are grammatical and make sense, except unlikely as translations.

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a mortal year

    And to begging again:

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> have a kindness in the world

    Again, all my commentary is meant as fun, I'm not intending to (mis)lead you to something here.

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a dead dog

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> put ā € |

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> have a deadline

    And more threats, again:

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a hash of you

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a mortal beefed up

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> have a heartbroker

    A heartbroker? Really? Something new.

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a hash of a tree

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> to put a lot of light on it

    And finally, the network gets hungry:

    ааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> to have a meal

    And positively concludes:

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a date auspicious

    аааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааааа -> a friend of a thousand years

    Hope you had fun reading these, and please try some for yourselves.

    Saturday, October 28, 2017

    What are some funny Google Translate tricks?

    This is the third quoranswer blog post, answering the question What are some funny Google Translate tricks? I have decided to update the Google translations based on the current situation. I think they are still a lot of fun. Let me know in comments, if you came across some funny translations!


    There used to be a funny politically coloured trick for Russian->English, where sense was inverted on translation depending on what President names were used in positive vs negative context. I can’t reproduce it right now, but GT produces this at the moment:
    Обама не при чём, виноват Путин.
    human: Obama is innocent, Putin is to blame.
    GT: Obama has nothing to do with Putin. (Previously in Aug 4, 2016: "Obama is not to blame, blame Putin.")
    Путин не при чём, виноват Обама
    human: Putin is innocent, Obama is to blame.
    GT: Putin has nothing to do with Obama's fault. (Previously in Aug 4, 2016: "Putin is not being Obama's fault.")

    Tuesday, October 24, 2017

    What grammatical challenges prevent Google Translate from being more effective?

    Here is one more Quora question on the exciting topic of machine translation and my answer to it.

    The question had some sub-questions:

    • Is there a set of broad grammatical rules which decreases its efficacy?
    • How can these challenges be overcome? Is it possible to fully automate good quality translation?

    Below is my answer, hoping it will be interesting to learn about machine translation and different language pairs. Note, that translations given currently by Google Translate might differ from below as they were obtained in 2013. UPD: and they do! See comments to this post.

    Google is pretty good at modeling close enough language pairs. By close enough I mean languages that share multiple vocabulary units, have similar word order, morphological richness level and other grammatical features.

    Let's pick an example of a pair, where Google Translate (GT) is good. Round-trip method is one way to verify whether the languages are close enough, at least statistically, for GT:

    (these examples are using GT only, no human interpretation involved)

    English: I am in a shop.
    Dutch: Ik ben in een winkel.
    back to English I'm in a store. (quite ok)

    English: I danced into the room.
    Dutch: Ik danste in de kamer.
    back to English: I danced in the room. (preposition issues)


    Let's pick a pair of more unrelated languages (by the way, when we claim the languages are unrelated grammatically, they may also be unrelated semantically or even pragmatically: different languages were created by people to suit their needs at particular moments of history). One such pair is English and Finnish:

    Finnish: Hän on kaupassa.
    English: He is in the shop.
    Finnish: Hän on myymälä. (roughly the original Finnish sentence)

    This example has pronoun hän, which in Finnish is not gender specific. It should be resolved based on larger context, than just a sentence. Somewhere before this sentence in a text, there should have been a mention of who hän is referring to.

    To conclude this particular example: Google Translate translates on a sentence level and that is a limitation in itself, that makes correct pronoun resolution impossible. Pronouns are useful, if we wanted to understand, what was the interaction between the objects in a text.


    Let's pick another example of unrelated languages: English and Russian.

    Russian: Маска бывает правдивее и выразительнее лица.
    English: The mask is truthful and expressive face. (should have been: The mask can be more truthful and expressive than face)
    back to Russian: Маска правдивым и выразительным лицом. (hard to translate, but the meaning roughly: The mask being a truthful and expressive face).

    To conclude this example: languges with rich morphology that, in the case of the Russian language, convey grammatical case in just a word inflection and thus require deeper grammatical analysis, which pure statistical machine translation methods lack no matter how much data has been acquired. There exist methods of combining rules and statistics together.


    Another pair and different example:
    English: Reporters said that IBM has bought Lotus.
    Japanese: 記者は、IBMがロータスを買っていると述べた。
    back to English: The reporter said that IBM Lotus are buying.

    Japanese has a "recursive syntax", that represents this English sentence, like:

    Reporters (IBM Lotus has bought) said that.

    i.e. the verb is syntacically placed after the subject-object pair of a sentence or a sub-sentence (direct / indirect object).

    To conclude this example: there should exist a method of mapping syntax structures as larger units of the language and that should be done in a more controlled fashion (i.e. is hard to derive from pure statistics).

    Saturday, September 23, 2017

    What's a good topic for a bachelor's thesis in Sentiment Analysis?

    Preamble

    Over the past few months (soon close to a year) you, my readers, might have noticed decline in frequency of my blogging. There are few reasons, including practical (absence of time), but still the most two important are:

    1. Blogger has not developed too much as a tool over time. It probably continues to be relatively popular and bringing some ad money, so Google did not shut it down. Moving over to medium.com might be a better idea in order to produce visually "shinier" posts and actually enjoy writing.

    2. There are other interesting and more interactive ways to share one's knowledge. One of such, that I personally like, is quora.com. The site offers a reverse model compared to blogging: you answer questions. This way you ensure, that at least the questioner will read your answer, but so might do other respondents. Rating of your answers is another component, that contributes to statistics and getting analogy of payment - credits, that you can later use for instance for boosting your answers to a larger audience. But I would say the latter is of lesser importance to me.

    Since I have never actually figured out, whether Quora allows you to read posts without being registered, re-posting my answers here from time to time could be a good way to also maintain this blog alive.

    So here we go (slightly edited version):

    What's a good topic for a bachelor's thesis in Sentiment Analysis?

    Apart from applying deep neural networks to sentiment analysis being exciting, another topic that is exciting both from research and practice perspective is sarcasm detection. It goes somewhat outside of the topic of sentiment analysis per se out to the opinion mining. Sentiment analysis precision and recall are affected by the sarcastic posts. This is because sarcastic posts tend to be positive on the surface (in fact to the conventional algorithms — ML based or rule-based ones), but suggest negative context.
    There are interesting situations that arise as a result of failing to recognize sarcasm. Borrowing from [1]:

    User 1 tweet:

    You are doing great! Who could predict heavy travel between #Thanksgiving and #NewYearsEve. And bad cold weather in Dec! Crazy!

    Response from a major U.S. Airline:

    We #love the kind words! Thanks so much.

    User 1:

    wow, just wow, I guess I should have #sarcasm

    User 2:

    Ahhh..**** reps. Just had a stellar experience w them at Westchester, NY last week. #CustomerSvcFail

    Response from a major U.S. Airline:

    Thanks for the shout-out Bonnie. We’re happy to hear you had a #stellar experience flying with us. Have a great day.

    User 2:

    You misinterpreted my dripping sarcasm. My experience at Westchester was 1 of the worst I’ve had with ****. And there are many.
    [1]
    Rajadesingan
    A. et al. Sarcasm Detection on Twitter: A Behavioral Modeling Approach Sarcasm Detection on Twitter

    Sunday, October 9, 2016

    Luke 6.2.1 release and all things open source

    Release

    Indeed, luke 6.2.1 for lucene 6.2.1 is out of the oven. This is the proud moment for Tomoko Uchida, my co-committer to have been a release manager for the first time. Congrats, Tomoko!

    Community

    As luke gets more and more stargazers on github (520 at the time of this writing), I tend to glance over the list of them which sometimes makes my day. But beyond that and more importantly, this lays out the community of Lucene / Solr / Elasticsearch users and developers, that hopefully enjoy using luke too. 

    Big names on user list

    Having access to the stats of the luke repo gives insights on who and when might be talking about luke. This time, it is PayPal Engineering. And here is their nice technical writeup on indexing lots of data in Elasticsearch and field usage of luke for optimizing the lucene index data structures: https://www.paypal-engineering.com/2016/08/10/powering-transactions-search-with-elastic-learnings-from-the-field/

    London Lucene/Solr hackday

    Hackday is an amazing way to jump out of a routine and think big: what can be improved in the search land of Lucene / Solr technology and tooling? It was great to see that luke was picked up as one topic on the Lucene / Solr hackday in London: https://github.com/flaxsearch/london-hackday-2016. And there it is, Marple, browser-driven explorer for lucene indexes: https://github.com/flaxsearch/marple. Go check it out.

    New contributors to luke

    Tomoko and I have been active promoting luke on various occasions, Lucene / Solr Revolution 2015 and  ApacheCon 2015. And of course on twitter. Recently Florian Hopf has become active in sending pull requests to improve luke and fix various nagging issues. Welcome!

    Wednesday, April 13, 2016

    Luke 6.0 has been released

    #luke 6.0 has been released. Major upgrade to #lucene 6.0 api: https://github.com/DmitryKey/luke/releases/tag/luke-6.0.0


    There are other interesting features cooking, like access to DocValues: https://github.com/DmitryKey/luke/pull/53

    If you feel like contributing, either by code or documentation, feel free to join the project:


    Wednesday, December 30, 2015

    Apache Solr Enterprise Search Server -- Third edition

    This year gave me a chance to be a technical reviewer of the book with search engine topic. The title is Apache Solr Enterprise Search Server and it saw the light in its third edition. The first edition back in 2010 helped me to start thinking in NoSQL way, despite that SQL has been literally everywhere (well, and still is). It does take a bit of mind warping to think beyond relational database lingo and data modelling and in my opinion is rather useful for your career as a software engineer.



    Here goes my review on Amazon:

    This book in its first edition was the first one around back in 2010, that covered Apache Solr in as much detail as I needed to get into the topic quickly. This third edition includes revisions for Apache Solr 5, notoriously covering things like Solr admin page, SolrCloud, scaling the search engine for large amount of documents, text analysis, indexing, search and even map-reducing your Solr index! In particular, throwing a MapReduce task at large-scale indexing task has been hard / unclear in the past and now it is available to any user of Apache Solr out of the box. This makes books like this immensely important to not waste one's time in looking around for useful bits of information scattered here and there. More importantly, authors of the book are directly involved into the project, either as Apache Solr / Lucene committers or active practitioners and developers of the technology. So I recommend this book for an entry-level and mid-level search engineers that look into getting their hands dirty with search problems and / or improving on the previously untapped areas of the search engine world.

    Sunday, October 11, 2015

    [ANNOUNCE] Luke 5.3.0 released: naturally runs on Java 8

    This release runs on Java8 and does not run on Java7.

    This release includes a number of pull requests and github issues. Worth mentioning:
    #38 upgrade to 5.3.0 itself
    #28 Added LUKE_PATH env variable to luke.sh
    #35 Added copy, cut, paste etc. shortcuts, using Mac command key
    #34 Fixed lastAnalyzer retrieval (this feature remembers the last used analyzer on the Search tab)
    #31 200 stargazers on github (by the time of this release the number crossed 260). Luke community is growing.

    Everybody is welcome to contribute. If you feel like you care about search / indexing or would like to get deeper with Apache Lucene, go ahead and pick a ticket: https://github.com/DmitryKey/luke/issues
    And, don't be afraid, we do not have any complaint departments:


    All you need is your favourite beverage and a good debugger.

    Wednesday, July 8, 2015

    [ANNOUNCE] Luke 5.2.0 released

    This is a major release supporting lucene / solr 5.2.0. Download the zip here:

    It supports elasticsearch 1.6.0 (lucene 4.10.4)
    Issues fixed:
    #20 Added support for reconstructing field values of indexed and not stored fields, that do not expose positions.
    Pull requests:
    #23 Elasticsearch support and Shade plugin for assembly
    #26 added .gitignore to project
    #27 Lucene 5x support
    #28 Added LUKE_PATH env variable to luke.sh
    #30 Luke 5.2

    I'd like to highlight the contribution of Tomoko Uchida who has been recently very active in sending pull requests, including upgrade to lucene 5.x and first version of Apache Pivot based luke ui.

    Wednesday, April 15, 2015

    Luke gets support for Elasticsearch indices

    That is that, really. The so long awaited proper support for elasticsearch indices.





    Luke supported Apache Solr indices already. Why not Elasticsearch? The reason was, that ES uses its own SPI for postings format. If you tried to open an Elasticsearch index with luke before, you'd get something like:

    A SPI class of type org.apache.lucene.codecs.PostingsFormat with name 'es090' does not exist. You need to add the corresponding JAR file supporting this SPI to your classpath. The current classpath supports the following names: [Lucene40, Lucene41]


    The biggest issue of supporting custom SPI is that you'd need to hack the luke jar binary and add the ES SPI. I bet it is not what you would want to spend your time on.

    With the excellent pull request by apakulov https://github.com/DmitryKey/luke/pull/23 luke uses shade maven plugin, that does all the magic. It magically updates the in-binary META-INF/services file with the following entry:

    org.elasticsearch.index.codec.postingsformat.Elasticsearch090PostingsFormat
    org.elasticsearch.search.suggest.completion.Completion090PostingsFormat
    org.elasticsearch.index.codec.postingsformat.BloomFilterPostingsFormat
    


    Currently this is available on luke master: https://github.com/DmitryKey/luke and a pre-release: https://github.com/DmitryKey/luke/releases/tag/luke-4.10.4-field-reconstruction

    Saturday, March 21, 2015

    Flexible run-time logging configuration in Apache Solr 4.10.x

    In a multi-shard setup it is useful to be able to change log level in runtime without going to each and every shard's admin page.

    For example, we can set the logging to WARN level during massive posting sessions and back to INFO, when serving the user queries.

    In solr 4.10.2 these one-liners do the trick:

    # set logging level to WARN,
    # saves disk space and speeds up massive posting 
    curl -s http://localhost:8983/solr/admin/info/logging \
                           --data-binary "set=root:WARN&wt=json" 
     
    # set logging level to INFO,
    # suitable for serving the user queries 
    curl -s http://localhost:8983/solr/admin/info/logging \
                           --data-binary "set=root:INFO&wt=json"
    

    Back from Solr you get a JSON with the current status of each configured logger.

    Monday, March 16, 2015

    Luke keeps getting updates and now on Apache Pivot

    Originally developed for fun and profit by Andrzej Bialecki, the lucene toolbox luke continues to be developed. Its releases are published at: https://github.com/DmitryKey/luke/releases


    Most recently Tomoko Uchida has contributed into effort of porting Luke to an Apache License 2.0 friendly GUI framework Apache Pivot. New branch has been created to host this work:

    https://github.com/DmitryKey/luke/tree/pivot-luke

    Currently supported Lucene: 4.10.4.

    It is far from completion, but already now you can:

    • open your Lucene index and check its metadata

    • page through the documents and analyze fields


    • search the index

    We will appreciate if you could test the pivot luke and give your feedback.

    Monday, November 17, 2014

    Lightweight Java Profiler and Interactive svg Flame Graphs

    A colleague of mine has just returned from the AWS re:Invent and brought in all the excitement about new AWS technologies. So I went on to watching the released videos of the talks. One of the first technical ones I have set on watching was Performance Tuning Amazon EC2 Instances by Brendan Gregg of Netflix. From Brendan's talk I have learnt about Lightweight Java Profiler (LJP) and visualizing stack traces with Flame Graphs.

    I'm quite 'obsessed' with monitoring and performance tuning based on it.
    Monitoring your applications is definitely the way to:

    1. Get numbers on performance inside your company, spread them and let people talk stories about them.
    2. Tune the system in where you see the bottleneck and measure again.

    In this post I would like to share a shell script that will produce a colourful and interactive flame graph out of a stack trace of your java application. This may be useful in a variety of ways, starting from an impressive graph for you slides to making informed tuning of your code / system.

    Components to build / install

    This was run on ubuntu 12.04 LTS.
    Checkout the Lightweight Java Profiler project source code and build it:

    svn checkout \
        http://lightweight-java-profiler.googlecode.com/svn/trunk/ \
        lightweight-java-profiler-read-only 
     
    cd lightweight-java-profiler-read-only/
    make BITS=64 all
    

    (omit the BITS parameter if you want to build for 32 bit platform).

    As a result of successful compilation you will have a liblagent.so binary that will be used to configure your java process.


    Next, clone the FlameGraph github repository:

    git clone https://github.com/brendangregg/FlameGraph.git

    You don't need to build anything, it is a collection of shell / perl scripts that will do the magic.

    Configuring the LJP agent on your java process

    Next step is to configure the LJP agent to report stats from your java process. I have picked a Solr instance running under jetty. Here is how I have configured it in my Solr startup script:

    java \
    -agentpath:/.../lightweight-java-profiler-read-only/\
          build-64/liblagent.so \
    -Dsolr.solr.home=cores start.jar

    Executing the script should start the Solr instance normally and will be logging stack trace to traces.txt

    Generating a Flame graph

    In order to produce a flame graph out of the LJP stack trace you will need to perform the following:

    1. Convert LJP stack trace into a collapsed form that FlameGraph understands.

    2. Call flamegraph.pl tool on the collapsed stack trace and produce the svg file.


    I have written a shell script that will do this for you.

    #!/bin/sh
    
    # change this variable to point to your FlameGraph directory
    FLAME_GRAPH_HOME=/home/dmitry/tools/FlameGraph
    
    LJP_TRACES_FILE=${1}
    FILENAME=$(basename $LJP_TRACES_FILE)
    
    JLP_TRACES_FILE_COLLAPSED=\
       $(dirname $LJP_TRACES_FILE)\
           /${FILENAME%.*}_collapsed.${FILENAME##*.}
    FLAME_GRAPH=\
           $(dirname $LJP_TRACES_FILE)/${FILENAME%.*}.svg
    
    # collapse the LJP stack trace
    $FLAME_GRAPH_HOME/stackcollapse-ljp.awk $LJP_TRACES_FILE > \
        $JLP_TRACES_FILE_COLLAPSED
    
    # create a flame graph
    $FLAME_GRAPH_HOME/flamegraph.pl $JLP_TRACES_FILE_COLLAPSED > \
        $FLAME_GRAPH
    


    And here is the flame graph of my Solr instance under the indexing load.



    You could interpret this diagram bottom-up: the lowest level is entry point class that starts the application. Then we see that CPU-wise two methods are taking the most of the time: org.eclipse.jetty.start.Main.main and java.lang.Thread.run.

    This svg diagram is in fact an interactive one: load it in the browser and click on the rectangles with methods you would like to explore more. I have clicked on the
    org.apache.solr.update.processor.UpdateRequestProcessor.processAdd rectangle and drilled down to it:


    It is this easy to setup a CPU performance check for your java program. Remember to monitor before tuning your code and wear a helmet.

    Friday, November 14, 2014

    Ruby pearls and gems for your daily routine coding tasks

    This is a list of ruby pearls and gems, that help me in my daily routine coding tasks.




    1. Retain only unique elements in an array:

    a = [1, 1, 2, 3, 4, 4, 5]
    
    a = a.uniq # => [1, 2, 3, 4, 5]
    

    2. Command line options parsing:

    require 'optparse'
    class Optparser
    
    def self.parse(args)
      options = {}
      OptionParser.new do |opts|
        opts.banner = "Usage: example.rb [options]"
    
        opts.on("-v", "--[no-]verbose", "Run verbosely") do |v|
         options[:verbose] = v
        end
    
       opts.on("-o", "--require OUTPUTDIR", "Output directory") do |o|
         options[:output_dir] = o
       end
    
       options[:source_dir] = []
         opts.on("-s", "--require SOURCEDIR", "Source directory") do |s|
         options[:source_dir] << s
       end
    
       end.parse!
    
       options
      end
    end
    
    options = Optparser.parse(ARGV) #pp options  When executed with -h, this script will automatically show the options and exit.  
    

    3. Delete a key-value pair in the hash map, where the key matches certain condition:

    hashMap.delete_if {|key, value| key == "someString" }
    

    Certainly, you can use regular expression based matching for the condition or a custom function, say, on the 'key' value.


    4. Interacting with mysql. I use mysql2 gem. Check out the documentation, it is pretty self-evident.

    5. Working with Apache SOLR: rsolr and rsolr-ext are invaluable here:

    require 'rsolr'
    require 'rsolr-ext'
    solrServer = RSolr::Ext.connect :url => $solrServerUrl, :read_timeout => $read_timeout, :open_timeout => $open_timeout
    
    doc = {field1=>"value1", "field2"=>"value2"}
    
    solrServer.add doc
    
    solrServer.commit(:commit_attributes => {:waitSearcher=>false, :softCommit=>false, :expungeDeletes=>true})
    solrServer.optimize(:optimize_attributes => {:maxSegments=>1}) # single segment as output
    

    Tuesday, September 23, 2014

    Indexing documents in Apache Solr using custom update chain and solrj api

    This post focuses on how to target custom update chain using solrj api and index your documents in Apache Solr. The reason for this post existence is because I have spent more than one hour figuring this out. This warrants a blog post (hopefully for other's benefit as well).

    Setup


    Suppose that you have a default update chain, that is executed in every day situations, i.e. for majority of input documents:

    <updaterequestprocessorchain default="true" name="everydaychain">
    <processor class="solr.LogUpdateProcessorFactory" />
    <processor class="solr.RunUpdateProcessorFactory" />
    </updaterequestprocessorchain>
    

    In some specific cases you would like to execute a slightly modified update chain, in this case with a factory that drops duplicate values from document fields. For that purpose you have configured a custom update chain:

    <updaterequestprocessorchain name="customchain">
    <processor class="solr.UniqFieldsUpdateProcessorFactory" >
    <lst name="fields">
       <str>field1</str>
    <lst>
    <processor class="solr.LogUpdateProcessorFactory" />
    <processor class="solr.RunUpdateProcessorFactory" />
    </updaterequestprocessorchain>
    

    Your update request handler looks like this:

    <requesthandler class="solr.UpdateRequestHandler" name="/update">
    <lst name="defaults">
    <str name="update.chain">everydaychain</str>
    </requesthandler>
    

    Every time you hit /update from your solrj backed code, you'll execute document indexing using the "everydaychain".

    Task


    Using solrj, index documents against the custom update chain.

    Solution


    First before diving into the solution, I'll show the code that you use for normal indexing process from java, i.e. with every:

    HttpSolrServer httpSolrServer = null;
    try {
         httpSolrServer = new HttpSolrServer("http://localhost:8983/solr/core0");
         SolrInputDocument sid = new SolrInputDocument();
         sid.addField("field1", "value1");
         httpSolrServer.add(sid);
    
         httpSolrServer.commit(); // hard commit; could be soft too
    } catch (Exception e) {
         if (httpSolrServer != null) {
             httpSolrServer.shutdown();
         }
    }
    

    So far so good. Next turning to indexing with custom update chain. This part of non-obvious from the point of view of solrj api design: having an instance of SolrInputDocument, how would one access a custom update chain? You may notice, how the update chain is defined in the update request handler of your solrconfig.xml. It uses the update.chain parameter name. Luckily, this is an http parameter, that can be supplied on the /update endpoint. Figuring this out via http client of the httpSolrServer object led to nowhere.

    Turns out, you can use UpdateRequest class instead. The object has got a nice setParam() method that lets you set a value for the update.chain parameter:

    HttpSolrServer httpSolrServer = null;
            try {
                httpSolrServer = new HttpSolrServer(updateURL);
    
                SolrInputDocument sid = new SolrInputDocument();
                // dummy field
                sid.addField("field1", "value1");
    
                UpdateRequest updateRequest = new UpdateRequest();
                updateRequest.setCommitWithin(2000);
                updateRequest.setParam("update.chain", "customchain");
                updateRequest.add(sid);
    
                UpdateResponse updateResponse = updateRequest.process(httpSolrServer);
                if (updateResponse.getStatus() == 200) {
                    log.info("Successfully added a document");
                } else {
                    log.info("Adding document failed, status code=" + updateResponse.getStatus());
                }
            } catch (Exception e) {
                e.printStackTrace();
                if (httpSolrServer != null) {
                    httpSolrServer.shutdown();
                    log.info("Released connection to the Solr server");
                }
    
            }
    

    Executing the second code will trigger the LogUpdateProcessor to output the following line in the solr logs:

    org.apache.solr.update.processor.LogUpdateProcessor  –
       [core0] webapp=/solr path=/update params={wt=javabin&
          version=2&update.chain=customchain}
    

    That's it for today. Happy indexing!