Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Sunday, December 15, 2019

Conferences in 2019

As Christmas and New Year holidays are coming up, I wanted to reflect on two conferences in two different countries -- Estonia and Ukraine, that I had a pleasure to participate and / or organize this year.

AINL

AINL Conference (https://ainlconf.ru/2019/program), held in Tartu in November, focused a lot on applying deep learning to NLProc, with two tutorials by Dmitry Ustalov (Yandex) on Crowdsourcing on Language Resources and Evaluation and by Andrey Kutuzov (University of Oslo) on Diachronic Word Embeddings for Semantic Shifts Modelling. Andrey Kutuzov's tutorial was practical and involved some Python coding, resulting in a pull request: https://github.com/wadimiusz/diachrony_for_russian/pull/5 that I submitted for the task of comparing semantic shifts in meaning between Soviet and Post Soviet eras. This code uses Jaccard similarity as a local method for detecting shifts in meaning. There are also global methods, like Procrustes alignment, the only downside of which is it is slower, than Jaccard. You can read more detail on the task in Andrey's AINL slides.


Credit: Dmitry Kan



In terms of submitted papers -- the review process was double-blind and involved at least 3 reviewers per paper. The result was 30% acceptance rate and 12 out of 40 papers that did make it, focused on data acquisition and annotation, human-computer interaction, statistical NLProc (including paper by Ansis Bērziņš on usage of speech recognition for determining language similarity -- video) and neural language models (one of the works for morpheme segmentation using Bi-LSTM model cited the work of Mathias Creutz with whom we worked at AlphaSense 2010-2016).

Last day of the conference focused on the industrial applications of AI in NLProc. By invitation of Lidia Pivovarova (University of Helsinki) I presented on the search engine and NLProc work we've done at AlphaSense, including smart synonyms, sentiment analysis, named entity recognition and salience resolution, theme modelling and high-precision search.

One of the challenges for the industrial presentation was that it had to last for 1,5 hours. If you consider your audience ability to focus only for 40 minutes, you have got to do something else than 65 slides. I decided to make about 30 slides and then handle the rest of my talk with Q&A. The outcome has been very surprising to myself, because the audience did want to learn details of AlphaSense product, making the Q&A last for 50 minutes. Quite a few questions I managed to answer with the product itself -- this sparks genuine interest in understanding the UI of an AI product powering the financial industry. I hope this was beneficial for the audience to dive into the workflows of financial knowledge workers and how NLProc can help solve their daily routine tasks better.


Customer Development Marathon

Customer development is the topic that interests me from the point of the product development. Just recently I've learnt about jobs-to-be-done approach to mining for real jobs that your customers hire your product for. One example with which Clayton Christensen of Harvard Business School motivates this approach is the job that male consumers of milkshakes had on the their way to work every day: stay engaged in life during monotonous driving and stay full until 10 a.m.

The conference (or marathon as we called it) on customer development attracted 70 participants at iHUB co-working center in Kyiv, Ukraine. Speakers from various established companies -- YouScan, MacPaw, PromoRepublic, Competera, AlphaSense, Kyivstar, Terrasoft, PMLab, Portmone.com, Weblium, VARUS, SendPulse, EVO.company -- presented 5 min talks about specific cases on engaging with their customers to grow conversion, retention and happiness with their products. Following the presentations, the discussion panels dug deeper into how to implement a customer-centric business. 


Credit: Maria Kudinova



We've organized the marathon in 3 panels: 

  1. Idea. Analysis. Validation 
  2. Creation. Delivery. Launch and
  3. Sales. Feedback. Innovation. 


Each of these panels focused on a particular stage of product development from idea to post-sale feedback and innovation loop. The audience learnt about how to conduct an efficient user interview, what tools help reach out to new or existing clients, how not to push your product into consulting or outsource, how to establish an internal company-wide communication to stay on the same page when shaping the product, marketing and sales around customer needs.

Both events were full of networking, meeting new and familiar faces in the industry and academia and learning a lot. For anything you aspire to build next year, focusing on real value and ease of use of your NLP / AI / search products, and thinking what job your users hire your products for will help you serve them better.

Tuesday, May 8, 2012

Paper on rule-based sentiment accepted!

My paper on rule-based sentiment was accepted to Dialog'2012, special section on ROMIP'2011. The ROMIP had a track on 2-way and 3-way sentiment classification of texts in Russian last year. In our team with @vporoshin we had three major systems:

1. Rule-based described in the paper.
2. Modified multinomial Naive Bayes trained on unigrams and bigrams.
3. Classifier ensemble of the two above.

Rule-based approach largely relies on the pre-crafted polarity dictionary. It means, that it knows only those polarity word sequences, that it has in the dictionary. The MNB classifier in contrast learns such sequences from training set. They also have other differences. MNB is in a way a bag-of-words approach, but may work surprisingly well. In 2-way classification it has shown accuracy of 90+% for one of the domains. The rule-based algorithm has interesting linguistic features, like object oriented sentiment detection. Although this first time, the ROMIP's sentiment tracks did not require an object oriented detection, the test data had an object name (e.g. movie title or product name) attributed to each text to classify. Both object oriented and general sentiment detection has performed equally well and above 50% (i.e. above the accuracy of a coin tossing method). Overall accuracy of the general rule-based classification is 63% with 92% precision for the positive class. This generally means that more polarity words should be mined for the negative class and the existing negative polarity dictionary revised (some words could be of positive or ambiguous polarity).

Some more numbers in the paper:

Sunday, March 18, 2012

Scientifc agenda of this year

This year stays promising in terms of the scientific happenings, first of all, I participated in the ROMIP contest on sentiment analysis. It was intense and interesting to dive into annotated and test data. More on this later, once information ready.

On the other note, this year's step up was to have been accepted on the committees list of the Second International Symposium on Business Modeling and Software Design (http://www.is-bmsd.org/). The research topics include and are not limited to the following:

BUSINESS MODELS AND REQUIREMENTS
- Business Analysis - Value Models and Process Models
- Essential Business Models
- Re-usable Business Models
- Relating Business Goals to Requirements
- Business Process Coordination
- Business Entities and Business Roles
- Business Data and Semantics
- Business Rules
- Behavior Modeling and Pragmatics
- Identification and Elicitation of Requirements
- Domain-imposed and User-defined Requirements
- Requirements Analysis

BUSINESS MODELS AND SERVICES
- Business Modeling and Service Science
- Relating Business Goals to the Identification of Services
- Service Modeling - Technology-independent and Platform-specific
- Business Rules and Service Composition
- Autonomic Service Behavior
- Context-aware Service Behavior
- Re-usable Service Models

BUSINESS MODELS AND SOFTWARE
- Business Modeling -driven Derivation of Software
- Business Innovation and Software Evolution
- Business-IT Alignment and Traceability
- Re-usable Business Models and Software Components
- Business Rules and Software Specification
- Business Goals and Software Integration
- Autonomic and Context-aware Business/Software Systems

INFORMATION SYSTEMS ARCHITECTURES
- Enterprise Architectures
- Service-Oriented Architectures
- Architectural Styles
- Architectural Viewpoints
- Crosscutting Concerns