Saturday, June 15, 2013

Solr on FitNesse

This year's Berlin Buzzwords conference was as intense as last year's. For me, in particular, it was heavier on the discussion side (hooked up with Robert Muir to discuss the "deduplication of postings lists" in Lucene and with Ted Dunning to speak some Russian), but some of talks have been interesting enough for me to try something practical immediately.

Dominik Benz of Inovex has presented on FitNesse tool.

In its own words: FitNesse is "the fully integrated standalone wiki and acceptance testing framework". Dominik was describing their experience with integrating it and told that the upfront investment is almost nil and suits to non-technical people. At this point I can confirm the former point, while the second needs more investigation really.

As the presentation concentrated quite heavily on how one would go about integrating FitNesse into the cycle of a Big Data project, I got curious whether this tool would be suitable for some of the tasks on Solr side. I have also compiled a presentation of my own, that summarizes what follows (some of the slides were borrowed from Dominik's slides).
A bit of thinking, and decided: implement a FitNesse fixture, that will check the health of solr cluster. Sometimes, when the cluster is too big (say, tens of nodes) someone could be overloading it with posting data or querying data. Some of the nodes (with solr shards) can go down or become unresponsive. It would be nice in a wiki setting to be able to say with a glance: is the cluster up and running or suffers for more CPU / RAM etc?

I'll present quite simple fixture for checking the solr health, which roughly took me 15 minutes to implement. I hope it can be useful for you too.

Here is how FitNesse UI looks like after executing the fixture:



The Java code:

package example;

import fit.ColumnFixture;
import org.apache.solr.client.solrj.SolrServer;
import org.apache.solr.client.solrj.SolrServerException;
import org.apache.solr.client.solrj.impl.CommonsHttpSolrServer;
import org.apache.solr.client.solrj.response.SolrPingResponse;

import java.io.IOException;
import java.net.MalformedURLException;

/**
 * Created with IntelliJ IDEA.
 * User: dmitry
 * Date: 6/14/13
 * Time: 3:48 PM
 * To change this template use File | Settings | File Templates.
 */
public class SolrShardsFixture extends ColumnFixture {

    private String shardURL;
    private String shardName;

    public boolean isShardUp() {
        if (shardURL == null || shardURL.isEmpty())
            throw new RuntimeException("shardURL url is empty");
        try {
            SolrServer serverTopic = new CommonsHttpSolrServer(shardURL);
            SolrPingResponse solrPingResponse = serverTopic.ping();

            if (solrPingResponse.getStatus() == 0)
                return true;

        } catch (MalformedURLException e) {
            throw new RuntimeException("Failed to create SolrServer instance: " + e.getMessage());
        } catch (IOException e) {
            throw new RuntimeException("Failed to ping the SolrServer instance: " + e.getMessage());
        } catch (SolrServerException e) {
            throw new RuntimeException(e.getMessage());
        }
        return false;
    }

    public void setShardURL(String shardURL) {
        this.shardURL = shardURL;
    }

    public void setShardName(String shardName) {
        this.shardName = shardName;
    }
}

Monday, May 6, 2013

MTEngine: switching UI languages

This is latest update from our test environment:

We have implemented a feature of respecting your browser language. That is, if you browser tells us en-us, we'll show the English version and if it is ru-ru, it is going to load the Russian version.



Friday, April 26, 2013

MTEngine: latest developments

Here are the latest developments going on on our test environment: MTEngine_test:

1. We took a snapshot of sentences from opencorpora.org and are working on pushing these into the new UI feature, called "tasks". Each task is one Russian sentence to be translated and rated by the user.
2. The feature with a free-form translation remains in the UI and is pushed into its own tab (screenshot in the Russian version of this message below).

For the production version of MTEngine we have done one improvement: when registering and using the system for the first time, the dictionary entries will be looked from the common dictionary, contributed by all our users.

Happy translations!


Same in Russian:

Свежие разработки в тест версии проекта MTEngine:

1. Мы взяли дамп предложений проекта opencorpora.org и работаем над новой фичей под названием "задания". Каждое задание -- это одно предложение на русском языке для перевода и оценки пользователем.
2. Фича с произвольным переводом пользовательских предложений на русском языке будет находится в отдельном табе:



Мы сделали улучшение и в продакшн версии: теперь, когда пользователь регистрируется и делает первые переводы, словарные единицы берутся из общего переводного словаря, который создали все пользователи проекта.

Успешных переводов и хороших выходных!


Friday, April 19, 2013

What grammatical challenges prevent Google Translate from being more effective?

Cross-posting my answer to the question in the topic on quora.com [1].

Google is pretty good at modeling close enough language pairs. By close enough I mean languages that share multiple vocabulary units, have similar word order, morphological richness level and other grammatical features.

Let's pick an example of a pair, where Google Translate (GT) is good. Round-trip method is one way to verify whether the languages are close enough, at least statistically, for GT:

(these examples are using GT only, no human interpretation involved)

English: I am in a shop.
Dutch: Ik ben in een winkel.
back to English I'm in a store. (quite ok)

English: I danced into the room.
Dutch: Ik danste in de kamer.
back to English: I danced in the room. (preposition issues)


Let's pick a pair of more unrelated languages (by the way, when we claim the languages are unrelated grammatically, they may also be unrelated semantically or even pragmatically: different languages were created by people to suit their needs at particular moments of history). One such pair is English and Finnish:

Finnish: Hän on kaupassa.
English: He is in the shop.
Finnish: Hän on myymälä. (roughly the original Finnish sentence)

This example has pronoun hän, which in Finnish is not gender specific. It should be resolved based on larger context, than just a sentence. Somewhere before this sentence in a text, there should have been a mention of who hän is referring to.

To conclude this particular example: Google Translate translates on a sentence level and that is a limitation in itself, that makes correct pronoun resolution impossible. Pronouns are useful, if we wanted to understand, what was the interaction between the objects in a text.


Let's pick another example of unrelated languages: English and Russian.

Russian: Маска бывает правдивее и выразительнее лица.
English: The mask is truthful and expressive face. (should have been: The mask can be more truthful and expressive than face)
back to Russian: Маска правдивым и выразительным лицом. (hard to translate, but the meaning roughly: The mask being a truthful and expressive face).

To conclude this example: languges with rich morphology that, in the case of the Russian language, convey grammatical case in just a word inflection and thus require deeper grammatical analysis, which pure statistical machine translation methods lack no matter how much data has been acquired. There exist methods of combining rules and statistics together.


Another pair and different example:
English: Reporters said that IBM has bought Lotus.
Japanese: 記者は、IBMがロータスを買っていると述べた。
back to English: The reporter said that IBM Lotus are buying.

Japanese has a "recursive syntax", that represents this English sentence, like:

Reporters (IBM Lotus has bought) said that.

i.e. the verb is syntacically placed after the subject-object pair of a sentence or a sub-sentence (direct / indirect object).

To conclude this example: there should exist a method of mapping syntax structures as larger units of the language and that should be done in a more controlled fashion (i.e. is hard to derive from pure statistics).


References
[1] http://www.quora.com/Linguistics/What-grammatical-challenges-prevent-Google-Translate-from-being-more-effective

Tuesday, March 12, 2013

MTEngine: new test UI features

MTEngine project is going forward and we are now testing new features in the test UI, that have been released last week:

1. Text boxes for editing the translation dictionary are wider now, comfortable for editing even on a mobile.
2. The translation progress indicator is now in the user focus area, right under the text area with the sentence in Russian.
3. The text area with the sentence in English has light-gray background now, similar to the one's of Google Translate service.

In the pipeline:
1. New translation history feature.
2. Linking VK profile for userpic in your user profile.
3. Content for pages "Download" and "About project" (in Russian only).

Feel free to join the testing! You need to know Russian and English.

Test UI URL:   http://semanticanalyzer.info/mtengine_test/

То же сообщение на русском:

Нововведения в тестовом UI:

1. Текстовые поля для правки словаря теперь шире, удобно даже на мобильном телефоне.
2. Индикатор прогресса перевода теперь находится в поле зрения пользователя, прямо под текстовым полем с предложением на русском языке.
3. Поле с переводом на английский теперь светло-серого цвета, как у Google Translate :)

На очереди:
1. Новая фича, где можно посмотреть историю своих переводов.
2. Подключение профиля вконтакте для загрузки юзерпика.
3. Контент для страниц "Скачать" и "О проекте".

Тестовый UI: http://semanticanalyzer.info/mtengine_test/

Thursday, December 6, 2012

Java Garbage Collector magic in action (or how to improve your java code using jconsole and jmap)

In my Java experience it has been somewhat unobvious how to jump from monitoring GC and fancy memory graphs with tools like jconsole to actually improving your code.

The ingredient I was missing apparently before was jmap that is part of JDK.

What the tool does is that it allows you to attach to a live java process by process id (pid) and output the histogram of its objects. Here is how it works:


jmap -histo 22170 > histo_22170.log

In the example, 22170 is the java process pid and command line option -histo makes jmap to output a histogram of objects. One nice thing about jmap is that it allows you build an object histogram on JVM OutOfMemory crash (see some details here).

The first lines of the histo_22170.log look like this, before some bug fixing has been done to the code (more on this in a moment):

 num     #instances         #bytes  class name
----------------------------------------------
   1:      88210219     2943436216  [B
   2:       5198407      455421864  [[B
   3:       5162015      123888360  com.mysql.jdbc.ByteArrayRow
   4:       2005702       95264296  [C
   5:       2006883       64220256  java.lang.String
   6:         37819       53525280  [I
   7:        309332       44543808  com.mysql.jdbc.Field
   8:        923280       36931200  java.util.TreeMap$Entry
   9:         18744       25864528  [Ljava.lang.Object;
  10:        471843       15098976  java.util.HashMap$Entry
  11:         36577        5195248  [Ljava.util.HashMap$Entry;
  12:         18196        4512608  com.mysql.jdbc.JDBC4PreparedStatement
  13:         18196        3202496  com.mysql.jdbc.JDBC4ResultSet
  14:         54308        2606784  java.util.TreeMap
  15:         12809        1903312  
  16:         36571        1755408  java.util.HashMap
  17:         12809        1750136  
  19:         18196        1164544  com.mysql.jdbc.PreparedStatement$ParseInfo

I have marked the relevant parts of the histogram with the bold font. The java process was doing some heavy-duty task for thousands of files and talking to the MySQL DB in a loop to load some meta-information for each of the file. The process was given 4GB max heap size and was not properly finishing, producing OutOfMemory error that in turn crashed the JVM.

The code snippet that was producing this looked like this:

PreparedStatement sqlStatement = sqlConnection.prepareStatement(
                          "SELECT * FROM SOME_TBL WHERE SOME_ID=?");
for(int i = 0; i < some_number_less_than_100; i++) {
    sqlStatement.setString(1, companyIds.get(i));
    ResultSet sqlResult = sqlStatement.executeQuery();
    if (sqlResult != null) {
     while (sqlResult.next()) {
        // do some processing of the query results here
     }
    }
}

Intuitively by now you should feel that something is wrong with the code around the JDBC object management.

Let's have a look on the bolded parts from the top of the object histogram. Apparently, the trending JDBC related objects are com.mysql.jdbc.Field with 309332 instances, com.mysql.jdbc.JDBC4PreparedStatement with 18196 instances and com.mysql.jdbc.JDBC4ResultSet with 18196 instances. Two latter objects have exactly same number of instances and that is reflected in our code, where both objects are re-created in a loop. The visual monitoring tool jconsole was showing constant RAM usage growth and Eden Heap Space being saturated with lots of young objects, while the Survivor Heap Space was not trending at all.

What's missing is releasing the JDBC resources, by calling close() methods on both PreparedStatement and ResultSet.

So let's correct the code:

PreparedStatement sqlStatement = sqlConnection.prepareStatement(
                           "SELECT * FROM SOME_TBL WHERE SOME_ID=?");
for(int i = 0; i < some_number_less_than_100; i++) {
    sqlStatement.setString(1, companyIds.get(i));
    ResultSet sqlResult = sqlStatement.executeQuery();
    if (sqlResult != null) {
     while (sqlResult.next()) {
        // do some processing of query results here
     }
     // missing lines added
     sqlResult.close();
}
sqlStatement.close();

After the two missing statements have been added (sqlResult.close() and sqlStatement.close()), the DB resources started to release properly and the original process began to work properly, without big spikes in RAM usage. The JDBC related objects have also disappeared from the top of the histogram:

 num     #instances         #bytes  class name
----------------------------------------------
   1:       1301417       66169256  [C
   2:       1333656       42676992  java.lang.String
   3:        158410       29981072  [I
   4:        382702       12246464  java.util.HashMap$Entry
   5:        116080        5668216  [B
   6:        230584        5534016  java.lang.StringBuffer
   7:         56760        3632640  java.util.regex.Matcher
   8:         18320        2620552 
   9:         18320        2501360 
  10:           588        2187960  [Ljava.util.HashMap$Entry;
  11:          1460        1733744 
  12:         33570        1523672 
  13:         19495        1247680  java.util.regex.Pattern
  14:         19729        1201136  [Ljava.lang.Object;
  15:          1460        1130688 
  16:         19477        1090712  [Ljava.util.regex.Pattern$GroupHead;
  17:          1312        1080992 
  18:         19096         614976  [Ljava.lang.String;
  19:         18642         596544  java.util.RandomAccessSubList
  20:         18642         596544  java.util.AbstractList$ListItr

Now the process is happily completing with reasonable RAM usage:


Click the image to make it bigger
The diagram shows that Eden Heap Space became much more free of young objects and the Survivor Heap Space gets utilized more. See here, if you want more details on various pools of Heap and Non-Heap memory.

Interestingly enough, this bug was hiding for months in the code base and only manifested itself once more data had to be processed. This made the process to run longer and thus reach and overflow the allocated RAM bounds.

This trivial example shows the importance of monitoring your heavy (and not so) java processes.

Happy monitoring!

Wednesday, June 6, 2012

Berlin buzz words 2012: impressions

This year I have had a unique chance to participate in the Berlin buzz words conference for the first time. In brief, it is the event where search, store and scale people come together to exchange on the recent ideas / developments in the area. I must say that the conference level simply amazed me: the quality of the presentations and the audience maturity have clearly aligned together.

Urania building, the venue


To me, as a Solr / Lucene user and developer it was especially fun to meet in person people I have previously only seen on the mail-lists or in video talks on the Internet. These, in particular, include (in my case): Otis Gostpodnetić, Uwe Schindler, Simon Willnauer, Robert Muir, Grant Ingersoll, Ted Dunning, Rafał Kuć. There've been new folks I haven't heard of previously and got inspired by their presentations, like Alex Lloyd from Google and Markus Weimer from Microsoft (opps, GOOG and MSFT in the same sentence). Got to see sematext guys in action at their SPM booth.


Opening session kicks in


The wi-fi worked everywhere, which is unnatural usually to other conferences. Yet, I kept my laptop at a hotel in order to force myself do three things: 1) actually listen to the presenter and ask questions via mike or in person; 2) occasionally take pictures; 3) network during the coffee-breaks.

First day's keynote session by Leslie Hawthorn


As a result: I took some amount of pictures; felt less distracted and tired at the end of each day; asked questions from the audience and got (probably) recorded on the video and many more questions in person; networked with leaders in their areas to actually perceive how things are going in their communities. SO this is to say, that in the end, what mattered to me was people and not only the technologies they have talked about.

Eric Evan's presentation


Some observations (probably interesting more to the conference orgs), pros and cons mixed:
1) The personal badge could have name on each side because of two reasons: it tends to always flip so that the name isn't visible and second - the map on the other side of it was useless, because it was easy to learn where each auditorium was.
2) Food was great and free beer / ice-cream / snacks by sponsors -- awesome addition.
3) Small auditoriums tended to have been super-packed and the only big one have been super sparse (excluding opening and closing sessions). Could be addressed somehow next year?
4) 20 minutes talks have been a surprise for the presenters that expected to have 40 min. The result is usually running out of time to ask any questions from the audience and presenters getting slowly to the core of their presentation.
5) Party on Monday evening and cute small surprises on the bus seats from wooga were cool!

There was also sometime left to explore the beautiful city of Berlin and of course eat Schnitzel!







Thanks to all the #bbuzz team for excellent experience and hoping to come next year!
yours truly,

Saturday, June 2, 2012

(first?) virtual presentation on Dialogue conference

Just participated in one of the biggest Russian conferences which fuses together theoretic and applied linguists, Dialogue'12. This time I couldn't come there in person, so instead we decided with @vporoshin to try out some modern technology. The selection was pretty easy: skype Finland->Russia, directed through speakers onto microphone connected to an amplifier. Also injected a photo of myself to add to "physical" presence. The conference organizers have appreciated utilizing new advanced technologies in presenting scientific papers. Here is the presentation (no author's photo there, you had to be present on the conference to see it):


Tuesday, May 8, 2012

Paper on rule-based sentiment accepted!

My paper on rule-based sentiment was accepted to Dialog'2012, special section on ROMIP'2011. The ROMIP had a track on 2-way and 3-way sentiment classification of texts in Russian last year. In our team with @vporoshin we had three major systems:

1. Rule-based described in the paper.
2. Modified multinomial Naive Bayes trained on unigrams and bigrams.
3. Classifier ensemble of the two above.

Rule-based approach largely relies on the pre-crafted polarity dictionary. It means, that it knows only those polarity word sequences, that it has in the dictionary. The MNB classifier in contrast learns such sequences from training set. They also have other differences. MNB is in a way a bag-of-words approach, but may work surprisingly well. In 2-way classification it has shown accuracy of 90+% for one of the domains. The rule-based algorithm has interesting linguistic features, like object oriented sentiment detection. Although this first time, the ROMIP's sentiment tracks did not require an object oriented detection, the test data had an object name (e.g. movie title or product name) attributed to each text to classify. Both object oriented and general sentiment detection has performed equally well and above 50% (i.e. above the accuracy of a coin tossing method). Overall accuracy of the general rule-based classification is 63% with 92% precision for the positive class. This generally means that more polarity words should be mined for the negative class and the existing negative polarity dictionary revised (some words could be of positive or ambiguous polarity).

Some more numbers in the paper:

Sunday, March 18, 2012

Scientifc agenda of this year

This year stays promising in terms of the scientific happenings, first of all, I participated in the ROMIP contest on sentiment analysis. It was intense and interesting to dive into annotated and test data. More on this later, once information ready.

On the other note, this year's step up was to have been accepted on the committees list of the Second International Symposium on Business Modeling and Software Design (http://www.is-bmsd.org/). The research topics include and are not limited to the following:

BUSINESS MODELS AND REQUIREMENTS
- Business Analysis - Value Models and Process Models
- Essential Business Models
- Re-usable Business Models
- Relating Business Goals to Requirements
- Business Process Coordination
- Business Entities and Business Roles
- Business Data and Semantics
- Business Rules
- Behavior Modeling and Pragmatics
- Identification and Elicitation of Requirements
- Domain-imposed and User-defined Requirements
- Requirements Analysis

BUSINESS MODELS AND SERVICES
- Business Modeling and Service Science
- Relating Business Goals to the Identification of Services
- Service Modeling - Technology-independent and Platform-specific
- Business Rules and Service Composition
- Autonomic Service Behavior
- Context-aware Service Behavior
- Re-usable Service Models

BUSINESS MODELS AND SOFTWARE
- Business Modeling -driven Derivation of Software
- Business Innovation and Software Evolution
- Business-IT Alignment and Traceability
- Re-usable Business Models and Software Components
- Business Rules and Software Specification
- Business Goals and Software Integration
- Autonomic and Context-aware Business/Software Systems

INFORMATION SYSTEMS ARCHITECTURES
- Enterprise Architectures
- Service-Oriented Architectures
- Architectural Styles
- Architectural Viewpoints
- Crosscutting Concerns

Monday, January 16, 2012

My experience with airBaltic

UPD: Please read the entire post. I will not remove the story line written originally, because this is exactly what has happened. However airBaltic contacted me on the phone themselves and told about positive resolution of the case. Please read on.

Original story:
-----
First of all, I would like to assure you that I'm not the best at blaming, meaning I simply don't like doing it publicly. It's probably unfair to only publicly blame an air operator and never praise them. But that's how it works. A happy customer doesn't compile an entire blog post about how cool it was to fly with a certain operator. "If I'm happy, I stay silent" principle. But believe me, if the case I'll tell you here about would resolve positively, I wouldn't hesitate to blog about it.

Here is the case. We planned a 3 days trip to Moscow from Helsinki and back together with my wife. Using skyscanner we've found the cheapest option: fly with airBaltic via Riga. Quick friend survey, all's good, settled. I went online and started my ticket search last Saturday. The cheapest option was to depart on 16:25. Chosen that, prepared to pay. After double-checking dates and times it struck me: nope, no good, departure set to 8:25. Cancelled search, started all over. First question here: is it bad luck or bad system? You choose.
Re-ran my search, all is good, paid 531 euros. But when I printed the travel receipt, this time it REALLY struck me: the return flight date was set to one month later! Another bad luck or system fault? This time I'm inclined to choose the second option. No problems, calling to the Finnish office. "On the weekends office is closed". I decided to call first thing next Monday morning. This is my first mistake and I admit it: should have called to Latvia and pay some euro and a half a minute to change the date.

Calling first thing Monday morning: young lady's voice, I described her the problem. She refused to change the date without an additional fee. I have asked her to connect me to her manager. After a couple of minutes (yeah-yeah, customer is on the first place), manager's voice: teaching and preaching me how I should have used their system. "On every page of multi-page ticket booking process, you can look on your right and check the departure and return flight dates and times." All right, thanks! But look, attempted I to explain the system fault: "Unless you really travel back in one month,there is no way to choose that different month without extra movements. No does the system suggest you the best return flights from in a month period!" This was simply noise for her and she continued teaching me how to use the system. I asked her to give me her manager / director. Guess what was the answer: "This is not possible". "Why?". "I'm sorry, but this is not possible." Being a customer, I'm pretty sure, I can talk to almost any worker of the company, who stays in the customer relations line. This time it is your fault, airBaltic.

"I would like to change the return flight date back to what it should be". But the lady tought me another time: "This is only possible if you pay 150 euros. You should have called us on Saturday and explained the problem." Which I did! And the Finnish office was closed. Is this really my problem now? I doubt it. Because, it is YOUR REPRESENTATIVE, airBaltic! What if I wouldn't have an opportunity to call abroad (yes, Riga is abroad to me) and pay for an international call? And if it didn't work, make sure to take my call on Monday morning seriously, attend to it and make an exception or a good men deal. What on Earth does this rule "if you called after two days, it cannot be changed without a fee" policy mean? Do you want to keep a customer or loose it? What do you loose by changing the month standing away date? Afraid not to find any cusomer during an entire month?

"And if I cancel the entire trip, what sum can be refunded?" "You get 76 euros back". Excellent.

Without further ramblings, I would like to publicly thank airBaltic for 531 euros worth "stay at home and don't travel with us" service. It has really taught me not to use your services. Ever.
Everything seems to be mortal in this wolrd, and airBaltic's serivce will die as well. But by making this type of "friendly" customer service and policies you only bring the end faster.

Good luck and enjoy 531-76 euros for not taking us where we wanted.
-----

UPDATE to the story: airBaltic continues working on the case, here is what they posted on twitter: @DmitryKan Dmitry, your case is not closed. Please give us a bit more time and colleagues will come back to you.

UPDATE 2: The case has been resolved. I have received a call from airBaltic, where they said that the return flight date was changed without an extra fee. I don't know was it a result of my social media activity since yesterday evening, but airBaltic service was extremely fast and accurate this time. Since all the posts I have done on the Internet about airBaltic link here, the landed people will read these updates as well. Thank you, airBaltic.

Wednesday, November 9, 2011

axis2: serialization and deserialization of wsdl2java generated objects

Using axis2's wsdl2java tool and a third-party wsdl I have generated service stub and supporting classes (data holders). Since there was a need to do post-processing of loaded data from a service, there was a need to serialize one of the data holder objects.

Questions that I had and posted on stackoverflow.com were:

1) is there a standard axis2 tool / approach that can be used for the purpose?

2) since the data holder class does not implement Serializable interface what would be the easiest way of serializing the object into xml format with the ability to restore the original object?

Data binding option was used (-d jaxbri) and each field of the class in question is annotated with @XmlElement tag, e.g.:

@XmlElement(name = "ID", required = true)
protected String id;



Here is how I solved it:

1. axis2 generated java classes set (client side) had an object called ObjectFactory. Majority of its methods create JAXBElement objects with values of fields of the class holder
2. I had to implement a serializable wrapper class ASerializable for the class holder, such that it uses the ObjectFactory to create the JAXBElement objects for all the fields.
3. some external code uses the wrapper class to create an serializable object and writes it to the output stream.
4. on the receiving end:

ASerializable aSerializable;
        A a;
        aSerializable= (ASerializable)in.readObject();
        a.setID((String)aSerializable.getID().getValue());

It still looks like extra work for the pre-annotated class serialization, but better than serializing into some text format and manual type checking during deserialization.
Some good intro into serialization with java can be found here.

Saturday, October 1, 2011

First international publication

Celebration!



Had my first international publication shown on the DBLP. Somehow it was something I wanted to achieve as an intermediate goal in the academic career. In a way this gives some visibility to what I have been doing for around 4 years. I mean NLP (Natural Language Processing) and Machine Translation more precisely. Before going international, I've had 5 publications in the Russian scientific journals and conferences.

On the same ICSOFT'11 conference where this publication has been presented in a form of poster, I had an honor to serve as knowledge-based systems track chair. Both presenting my work and leading the session were exciting. I wanted also to say thanks to the ICSOFT's organizing committee for giving me the participant's grant, that made my participation possible. Special thanks to Sergio Brissos.


ICSOFT 2011 Conference


ICSOFT is strictly not an NLP conference. However, it has a knowledge-based track, where rather relevant NLP related topics are listed:

Ontology Engineering
Decision Support Systems
Intelligent Problem Solving
Expert Systems
Reasoning Techniques
Knowledge Acquisition
Knowledge Mining
Machine Learning
Natural Language Processing
Human-Machine Cooperation

Two publications I remembered



From these, ontology engineering articles were strong. One of them (Barbara Furletti, Franco Turini: Mining Influence Rules out of Ontologies, see here) was about mining ontology and reasoning rules from the oldest Italian bank's data. This sounds exceptional to me, when some (even old) commercial data is given away to researchers.

Another, non-directly related to NLP, article I remembered was by Manolya Kavakli et al (Manolya Kavakli, Tarashankar Rudra, Manning Li: An Embodied Conversational Agent for Counselling Aborigines - Mr. Warnanggal.), where one of the challenges is providing health assistance to the Australlian aborigines via a computer based system, not very motivated people, poor, stealing food and other things. Another challenge is dealing with about 500 languages, that these aborigines speak. Here is a potential for interesting NLP problems.

Why would I recommend going to a conference not directly related to your research topic?


As a pre-word, I should mention, that in a way whatever we do in the NLP is materialized in the form of programming code. Therefore our work qualifies to a software engineering conference as well as to an NLP one.

Going to a strictly SW conference can give you the following benefits:
* concrete questions of you work in the light of software development practices. Some NLP researchers may think it is not very important to make their SW configurable, re-usable or performant. In the end of the day, this matters a lot, especially if you plan to implement you work into industrial level solution

* if you do a poster presentation, people can give you good insights into the quality of your poster and what can be improved. There were two extreme cases on the conference: one with the entire article text being pasted into the poster and another one with a couple of boxes and an arrow between them. The audience has reacted in an expected way: the first poster did not draw almost any attention, while the second had gathered the majority of the audience.

* you can pause and reflect a little bit: are you doing something valuable? Do you like what you do?


A couple of words about Spain, where ICSOFT'11 happened. +45 is something I have experienced for the first time; visiting royal palace Alcázar of Seville was extremely interesting and of course partying with conference peers over Spanish wine and tapas made the event memorable.

Enjoy you research life and publish your work as soon as possible.

Saturday, August 27, 2011

Оценка системы машинного перевода

Есть система машинного перевода с русского на английский.
Нужно сделать ручную оценку работы системы.

Целевая аудитория: все, кто хочет сделать машинные переводчики лучше (примеры: translate.google.com, translate.ru) и люди, интересующиеся прикладной лингвистикой (Natural Language Processing). Умение программировать НЕ требуется.

Связь: dmitry.kan[+AT+]gmail.com, twitter: DmitryKan

Задача: получить у меня пакет предложений (объём: сколько возьмётесь).

В пакете: предложения на русском языке и их переводы экспертом на английский язык.

Прогнать предложения на русском через систему. Просмотреть вручную их переводы на английский язык.

Составить список слов, которые не были найдены (это просто: не найденные слова будут выведены на русском). Послать мне список слов, я добавлю их в систему машинного перевода.

На выходе от вас три группы предложений из пакета:
1. хорошо перевелись
2. приемлемо перевелись (понятно по английской фразе, что было в русской)
3. плохо перевелись (непонятно по английской фразе, что было в русской)

Работа волонтёрская. Начальный бонус: статья со мной в соавторстве на конференции либо в журнале, если вам это интересно. Если нет -- всяческий пиар вам.

Дальше: если сработаемся, предложу Вам работу в лингвистических проектах (умение программировать обязательно).

Ссылки для интересующихся
[1] http://www.slideshare.net/dmitrykan/icsoft-2011-51cr
[2] http://www.slideshare.net/dmitrykan/automatic-build-of-semantic-translational-dictionary
[3] http://ufal.mff.cuni.cz/umc/

interested in rule based machine translation (rbmt)? / Интересуетесь машинным переводом на правилах?

I'm looking for students and activists of rule-based machine translation to help me in the evaluation of my machine translation system from Russian into English. Details in the e-mail: dmitry.kan[+AT+]gmail.com (substitute characters from [ fro ] with @).

Я ищу студентов и активистов машинного перевода на правилах для оценки моей системы машинного перевода с русского на английский. Детали по почте: dmitry.kan[+AT+]gmail.com (замените символы с [ по ] знаком @).

Thursday, July 14, 2011

Interested in machine translation between Russian and English?

Then mark August 15-19 2011 in your calendars. Web of Data'11 has accepted my poster on machine translation with semantic features, the full paper title is:

Semantic Feature Machine Translation System for Information Retrieval

Some details on the work from another poster, accepted to ICSOFT'11 can be checked here:

Sunday, July 10, 2011

Пример кода, интегрирующего AOT морфологический лемматайзер в C#

АОТ предлагает свой лемматайзер для русского и английского языка на сайте www.aot.ru. Если Вам нужно интегрировать их COM внутри проекта на C#, читайте ниже.

После установки библиотеки при помощи Setup.exe, загрузите Lemmatizer.dll в C# проект. Скопируйте следующий метод или его тело, например, в main-class:


private static void initAOTMorphoanalyzer()
{
LEMMATIZERLib.ILemmatizer lemmatizerRu = new LEMMATIZERLib.LemmatizerRussian();
lemmatizerRu.LoadDictionariesRegistry();
LEMMATIZERLib.IParadigmCollection piParadigmCollection = lemmatizerRu.CreateParadigmCollectionFromForm("мыла", 0, 0);

Console.Out.WriteLine(piParadigmCollection.Count);

for (int j=0; j < piParadigmCollection.Count; j++)
{
object[] args = { j };

Type paradigmCollectionType = piParadigmCollection.GetType();

if (paradigmCollectionType != null)
{
object Item = paradigmCollectionType.InvokeMember("Item", BindingFlags.GetProperty, null, piParadigmCollection, args);
Type itemType = Item.GetType();
if (itemType != null)
{
object Norm = itemType.InvokeMember("Norm", BindingFlags.GetProperty, null, Item, null);
Console.Out.WriteLine(Norm);
}
else
Console.Out.WriteLine("itemType is null");
}
else
Console.Out.WriteLine("paradigmCollectionType is null");
}
}



Результат:
2
МЫЛО
МЫТЬ

COM test example for C#: AOT lemmatizer

I will post this both in English and Russian for more people's benefit.

There is a Russian / English lemmatizer from AOT (www.aot.ru). If you need to use the COM that AOT provides inside C#, read on. Load the lemmatizer.dll inside your C# project. Insert the following method or its body inside your code, for example main class:


private static void initAOTMorphoanalyzer()
{
LEMMATIZERLib.ILemmatizer lemmatizerRu = new LEMMATIZERLib.LemmatizerRussian();
lemmatizerRu.LoadDictionariesRegistry();
LEMMATIZERLib.IParadigmCollection piParadigmCollection = lemmatizerRu.CreateParadigmCollectionFromForm("мыла", 0, 0);

Console.Out.WriteLine(piParadigmCollection.Count);

for (int j=0; j < piParadigmCollection.Count; j++)
{
object[] args = { j };

Type paradigmCollectionType = piParadigmCollection.GetType();

if (paradigmCollectionType != null)
{
object Item = paradigmCollectionType.InvokeMember("Item", BindingFlags.GetProperty, null, piParadigmCollection, args);
Type itemType = Item.GetType();
if (itemType != null)
{
object Norm = itemType.InvokeMember("Norm", BindingFlags.GetProperty, null, Item, null);
Console.Out.WriteLine(Norm);
}
else
Console.Out.WriteLine("itemType is null");
}
else
Console.Out.WriteLine("paradigmCollectionType is null");
}
}



Output:
2
МЫЛО
МЫТЬ

Wednesday, June 8, 2011

Amazed by Scala #1: objects and compilation

7 minutes and here is an object, that can be compiled into java classes:



import scala.actors._
import Actor._

object TopStock {
val symbols = List( "AAPL", "GOOG", "IBM", "MSFT")
val receiver = self
val year = 2008

def main(args: Array[String]) = {
symbols.foreach { symbol =>
actor { receiver ! getYearEndClosing(symbol, year) }
}

val (topStock, highestPrice) = getTopStock(symbols.length)
printf("Top stock of %d is %s closing at price %f\n", year, topStock, highestPrice)
}

def getYearEndClosing(symbol : String, year : Int) = {
val url = "http://ichart.finance.yahoo.com/table.csv?s="+
symbol + "&a=11&b=01&c=" + year + "&d=11&e=31&f=" + year+
"&g=m"

val data = io.Source.fromURL(url).mkString
val price = data.split("\n")(1).split(",")(4).toDouble
(symbol, price)
}

def getTopStock(count : Int) : (String, Double) = {
(1 to count).foldLeft("", 0.0) { (previousHigh, index) =>
receiveWithin(10000) {
case (symbol : String, price : Double) =>
if (price > previousHigh._2) (symbol, price) else previousHigh
}
}
}
}


Saved in TopStock.scala. Compiled with

> scalac TopStock.scala

ran with

> scala TopStock

Top stock of 2008 is GOOG closing at price 307,650000

Amazed by Scala

Top stock of 2008 is GOOG closing at price 307,650000 among (AAPL, GOOG, IBM, MSFT). Amazed by simplicity, clarity and beauty of the following code in Scala from this book.



import scala.actors._
import Actor._

val symbols = List( "AAPL", "GOOG", "IBM", "MSFT")
val receiver = self
val year = 2008

symbols.foreach { symbol =>
actor { receiver ! getYearEndClosing(symbol, year) }
}

val (topStock, highestPrice) = getTopStock(symbols.length)

printf("Top stock of %d is %s closing at price %f\n", year,
topStock, highestPrice)

def getYearEndClosing(symbol : String, year : Int) = {
val url = "http://ichart.finance.yahoo.com/table.csv?s="+
symbol + "&a=11&b=01&c=" + year + "&d=11&e=31&f=" + year+
"&g=m"

val data = io.Source.fromURL(url).mkString
val price = data.split("\n")(1).split(",")(4).toDouble
(symbol, price)
}

def getTopStock(count : Int) : (String, Double) = {
(1 to count).foldLeft("", 0.0) { (previousHigh, index) =>
receiveWithin(10000) {
case (symbol : String, price : Double) =>
if (price > previousHigh._2) (symbol, price) else
previousHigh
}
}
}