Wednesday, 25 April 2012

Computer language mystery solved by humans


Computers have languages, too. According to an article in the American Scientist, even the experts do not agree how many programming languages there are – estimates range from 2,500 to over 8,500.

One recent example which highlighted this variety was the mystery of the programming language used in the creation of “Duqu”, a computer Trojan which has been studied by heavyweight anti-virus companies like Symantec, Kaspersky Labs and F-Secure. These IT giants were able to see the code which this Trojan consisted of, but they were not able to identify which programming language had been used to compile this code.

Why didn’t they ask a computer?
To me, as a mere computer user without a programming background, the solution appears simple. It is a computer language, and a computer is obviously able to follow the instructions in the code (otherwise the Trojan would be of no use to the crooks who created it). So a computer should be able to identify what language it is. This seems to be an obvious logical conclusion.

But it is not so. Igor Soumenkov, a Kaspersky Lab Expert, wrote a blog article “The Mystery of the Duqu Framework”. The article outlines the history of the study of Duqu and the structure of the threat which it poses, and it ends with an appeal which amazed me: We would like to make an appeal to the programming community and ask anyone who recognizes the framework, toolkit or the programming language that can generate similar code constructions, to contact us or drop us a comment in this blogpost.

Digital guesswork?
Soumenkov received a flood of blog comments and e-mail responses, and the mystery of the programming language has now been solved. But it is interesting to check out the wording of the 159 comments on the original blog article. They are peppered with phrases like:
That code looks familiar
It may be a tool developed by ...
I think it's a ...
What about ...?
Just a guess ... the first thing that pops to my mind is ...
Sounds a lot like ...
I am not a specialist but I would say it could be ...
One more guess ...
This does smell to me a little bit like ...
I'm gonna take a wild guess ...
Plus a generous sprinkling of words like might, perhaps, maybe, probably, similar, clue, feel, remember, possibility and similar vague terms.

Data or brains?
For me, this throws an interesting light on the use of computers in natural language processing. The human guesswork in the comments on Duqu included many ideas that turned out to be wrong, but the brainstorming process was helpful to the computer experts involved, and the fuzzy process of human thinking led to a solution which evidently was not possible with the computer alone. And all of this for a language which is only useful in computers and has no meaning for human communication (when did you last _class_2.setup_class13)[esi]?).

The situation in translation between human languages is comparable. Automatic translation programs from Google, Microsoft, IBM and others can achieve a certain amount of pattern recognition and sometimes come up with plausible solutions. But only a competent human being can evaluate whether this solution is really accurate or appropriate. So these programs can be a useful tool in the hands of an expert, but there is a distinct risk that they may get the wrong end of the stick.

Friday, 2 March 2012

Would I advise my grandchildren to translate?

Bang, bang, bang.
Is this another nail in the coffin of freelance translation as a career?
A recent article on the blog of the Translation Automation User Society (TAUS) does not hold out much hope for specialist translators. The title of the article is “Who gets paid for translation in 2020?”. I would love to quote the author of this article by name, but no name is given. Perhaps this is a model article, generated by a computer, untouched by human hand. This would graphically illustrate the creed which underlies the article:
“In 2020 words are ‘free’. Almost every word has already been translated before. Our words will be stored somewhere and used again, legitimately in the eyes of the law or not. .... Even today ‘robots’ are crawling websites to retrieve billions of words that help to train machine translation engines. The latent demand for translation created by unprecedented globalization is making piracy an act of common sense.”
The TAUS vision paints a glowing picture of a completely automated future, with instant computerised translation in every hand-held device, every computer application and on every website, without any need for specialist intervention. To achieve this, TAUS aims to build up a database of all the translation work done in the world. It seems to envisage three methods to do this:
LinkBEG, SCAVENGE and STEAL
BEG: In conference lectures, blog articles and other publications, TAUS calls on translators to donate their translations to its central database. The reward for doing this is to know that we are contributing to the BRAVE NEW WORLD of global computerised translation. There may be some payback in the form of access to databases provided by others, but the rhetoric of the begging prose is that we should contribute for free to the ideal of a humanity without language barriers.
SCAVENGE: The above quote speaks of the “robots” which are retrieving billions of translated words to train machine translation engines. But a scavenger takes everything that it can find. A scavenger cannot afford to be fussy about quality. There are two experts in the industry who have important things to say about this. First of all Kirti Vashee in his blog eMpTy Pages. Kirti is an ardent advocate of machine translation, but he insists that the data used to train the translation engines must be of extremely high quality. The danger of the TAUS vision of innumerable robots scavenging for more and more data is that this can include lots of low quality data, so the resulting translations will be inherently problematical. The other expert is Miguel Llorens, a highly insightful freelance translator who ridicules many of the assumptions of the machine translation gurus and elegantly criticises buzzwords such as the “content tsunami” and “crowdsourcing”.
As an aside: Kirti and Miguel disagree on many things - I suppose it is not often that they are recommended as two leading experts in the debate on machine translation.
STEAL: It has often been suggested that Internet giants such as Google and Facebook are in fact data-gobbling monsters which think nothing of violating data protection standards. But at least in their public statements, they usually claim to respect the privacy of their users and to comply with data protection laws. Not so TAUS. In the above quotation, TAUS explicitly suggests that piracy is “an act of common sense”. I wonder if the similarity to the confiscation of private assets in the ideology of Marx, Stalin and others is merely accidental. Brave new world indeed!
Translation and my grandchildren
By the time the brave new world predicted by TAUS comes to pass (2020), my own translation career will be drawing to a close, or perhaps already ended. But what about my wonderful grandchildren? They will be on the threshold of their working lives (and some will be still in primary school). What should I tell them if they ask about translation as a career?
I will say: “Why not - if that is what you are really good at.” Of course I will point out the general principles of working in a career like translation: real language expertise in two languages, realistic self-appraisal and self-management, translating skills, the need for solid specialisation, how to use the tools of the trade (including computer-aided translation and various forms of machine translation), how to advertise and find customers and much more.
This is because essentially I do not accept the TAUS creed that “Almost every word has already been translated before.”. Even at the word level, in my work I regularly come across newly created terms or compound words (German legal and architectural prose has an amazing level of inventiveness in this respect). And at the sentence level, every language on earth has an incredible potential for creative new combinations of ideas and even new linguistic structures - after all, I believe that we are still building the tower and city of Babel.

Tuesday, 17 January 2012

12 facts, hints and ideas on databases in DVX2

Déjà Vu X2 is a “Translation Memory” program (TM). It does not come with pre-packaged language content. Instead, it remembers your own work, i.e. it acts as a “memory” for what you have “already seen” (= “déjà vu” in French).

1. There are three types of memory:
The TM (Translation Memory), the TB (Termbase) and the lexicon for each project.
  • The TM is a database where you can save the sentences from your source text together with your finished translation.
  • The TB is a terminology database which you can use for single words or whole phrases.
  • The lexicon is a database which only applies to the individual project. For every project file you can create a new lexicon.
When you then work on your project DVX2 combines the content of these three database types to suggest translations and help you in your work. The methods which DVX2 uses to make these suggestions are known as “Pretranslate”, “Assemble” and “AutoAssemble” – but that is another topic for another day.

2. Big Mama and Big Papa:
You can keep all of your work in just one TM (“Big Mama”) and one TB (“Big Papa”). If you are careful to give your entries the appropriate subject and client codes, DVX2 will take these codes into account when suggesting translations from your databases. My main TM contains about 40,000 sentence pairs accumulated over 12 years, and my main TB has about 55,000 entries.

3. Separate TMs and TBs:
In DVX2 Professional you can have up to 5 TMs and 5 TBs open in any project, and DVX2 Workgroup has no limitation. So you can use your Big Mama/Papa together with external databases, e.g. a TM or terminology list provided by the client, general reference material such as the EU DGT database, or terminology lists from major enterprises such as Microsoft, SAP or from various banks. Or you may even decide to keep separate databases for different subjects or clients instead of a Big Mama or Big Papa. You may feel that this is safer if you work on texts for competing engineering or IT firms which deliberately use different terminology for their own brands. The problem is that it may be more difficult to access all of your reference material, for example if you know that you have dealt with a term or sentence in DVX2, but you can’t remember which database you were using at the time.

4. Fuzzy matching:
You can allow DVX2 to find matching material which is not quite exact. Under Tools>Options>General you can set a percentage figure for the variants which DVX2 is allowed to find (= “Minimum Score”). The default setting is 75%, but depending on the type of inflections which occur in your languages it may be useful to set it to 50% or less. The percentage applies to both the TM and TB. It does not apply to the lexicon – only exact matches are found in the lexicon. And the “minimum score” does not affect the performance of the DVX2 functions DeepMiner and AutoWrite.


5. Adding new entries:
This is very quick and easy in DVX2. For the TM you enable AutoSend (either with the tick box at Tools>Options>Environment, or via the icons at the bottom of the DVX2 window – AutoSend is the second icon from the right). Then all you need to do is click CTRL-DownArrow when you have finished each segment. For the lexicon you have to highlight the word or phrase in the source and target text, then hit the F10 key. For the TB you again highlight the word or phrase in the source and target text, then hit F11. This brings up the following window:

Here you can edit the term in either language to add or remove declensions, correct spelling problems etc. You can check that the terms are marked with the right subject and client codes. There are additional fields, too (Definition, Part of Speech, Gender, Number, and you may also see a field called Context). I have not yet seen any reason to use any of these fields, although some users may have found ways to do so.
The termbase (TB) is one of the keys to productivity in DVX2. It is advisable to add words, and even whole phrases, as often as you can. Some users have the principle of adding an entry to the TB in every single sentence they translate. Steven Marzuola’s article about using the terminology database was based on the previous version of DVX (now often called DVX1), but it offers great advice which is also relevant to DVX2.

6. Subject and client codes:
These are important, because DVX2 refers to them when it decides what material to offer to help you with your current translation. When you first install DVX2, you will see a suggested list of subjects, but you can easily delete this and create your own list if you think this is better for your work. Each subject consists of a short index code (435 in my example above) and a descriptive text (Regional planning/ecology). When DVX2 decides how close the subject is to your current project, it works hierarchically, so in this example it would consider that entries with my subject codes 43 (Urban planning) and 4 (Building) are closely related. You can use letters instead of numbers if this suits your work.


7. Build lexicon:
This is a function which you can find in the “Lexicon” menu, and which is sometimes useful in preparation for a job which is heavy on terminology. I use this function for between 5% and 10% of my jobs. My procedure is as follows. First I call up “Build lexicon” and define the maximum number of words (usually 4). The program then takes a couple of minutes to find solutions. Then I open the lexicon (with the Project Explorer), click on the heading over the left hand column and define the sort criteria: 1. Number of words (descending), 2. Frequency (descending). Then I go through the list manually from the top. First I decide which four-word phrases are worth adding a lexicon entry for. This is usually only worthwhile for phrases which are meaningful in themselves and which occur frequently. When I get down to phrases which appear three times or less, I then use the scroll bar to move down to the most frequent three-word phrases. And so on, until I have defined a number of lexicon entries. Then I select “Remove entries” from the Lexicon menu, click on “Entries with empty targets” and OK. Typically, this gives me between 30 and 50 lexicon entries for a job consisting of several hundred segments, but they are entries which occur frequently and require consistency, so this preliminary process improves the results achieved by Pretranslate or Assemble as I work on the job.

This function (Build lexicon) can also be used to identify terms that can be used for a terminology list to be delivered to the client if this is part of the client’s instructions for the job. Over the years I have only had one such project, but this may be relevant for translators who often work in highly technical fields.

8. Names, places and proprietary titles:
These are the classic elements which should be added to the lexicon. If you have a product name or number, this is normally only relevant to the job in hand. You do not usually want this term to occur in jobs for other clients. The same applies to the names of the people who work for the client. Therefore, such elements should only be sent to the lexicon, and not to the termbase. But some names occur so often that they may be useful in the TB. My general principle here: if names could be confused with actual words in the language, they are not suitable for the TB. So the common German name Helmut is not in my TB because, depending on the level of fuzzy matching, it could be confused with the word Helm=helmet (and the declined forms Helme/Helmen/Helmes). Similarly, the surname Kohl is not in the TB to avoid confusion with Kohl=cabbage (and the near-match Kohle=coal). But the two names together are in the TB – i.e. the former German Chancellor Helmut Kohl. And other famous politicians are there too with the spelling in German and English, such as Gorbatschow/Gorbachev.
9. Adapting your use of the databases to your languages:
In some cases, your language pair and translation direction will influence the way you use the different databases because of issues such as word order and inflection. One example of this is the English phrase “public green spaces”. In French the words come in a different order, e.g. “espaces verts publics”, and alternative wordings are possible, e.g. “espaces verts des lieux publics”, “espace verts ouverts au public”, “espaces verts pour le public” etc. (Thanks to Dave Turner for providing these and other examples). In German the first translation that comes to mind is “öffentliche Grünflächen”, although the first word could also be declined as “öffentlichen”.

If you are translating from French to English, you will probably want to enter each and every French phrase as a lexical unit, especially if it occurs frequently in the type of text you deal with. Merely entering the elements does not help very much, because the order of the words must be changed. Depending on your type of work and the frequency of such phrases, you may decide to store them in the lexicon, the TB or the TM.
If you are translating from German, in this case it is sufficient to add the two words to the termbase and let DVX2 handle the endings as “fuzzy matches”. Even if we consider phrases with a greater number of inflected variants such as “public building”, (“öffentliche Gebäude”, “öffentliches Gebäude”, “öffentlichen Gebäudes”, “öffentlichem Gebäude”), it is still possible to enter just one version of each word and use fuzzy matching. The advantage here is that although the German source is inflected, the English target phrase is not.
Translating from a largely uninflected language into inflected languages like French and German can be more complicated, so you will have to find a strategy which fits the languages that you work with. There is no single solution which will work for all languages and all subject areas, but DVX2 offers flexibility in the use of the databases.

10. Looking things up in the database:
There are various ways to access the information that is in your databases. The first is that DVX2 uses this information to compile its suggested translation (when you use the functions “Pretranslate”, “Assemble” or “AutoAssemble”). When you have done that, you will see that some words or phrases in the suggested translation are underlined in blue. These are terms for which your databases contain several possibilities. Right clicking on the word or phrase will show you the other suggestions, and you can examine these and select them with the mouse or by using the number shown. The third way to see the relevant content of your database is by looking at the “Portions” window or windows. There are several screenshots illustrating this here. The fourth way to look up the information is to use Scan (CTRL-S) to call up a concordance from the TM, or Lookup (CTRL-L) to see entries from the TB.


11. Moving databases to another computer:
If you need to move your work to a different computer, e.g. to work on a laptop while you are travelling, you will need to copy certain files to the other computer. The first file is your project file, which has the extension .dvprj. The project file contains the lexicon, so no special steps are needed to transfer the lexicon. The termbase is a single file with the extension .dvtdb. The TM consists of at least four files. The main content is in a file with the extension .dvmdb. Then there is an index file for each of your languages; my index files have the extension en.dvmdi and de.dvmdi (for English and German). There is also a file with the extension .dvmdx. When you open the project on the other computer, DVX2 may complain that it cannot find the databases. But this is not a problem – when the project is open, you can select them with Project>Properties>Databases.

Another file which is worth moving to the other computer is the settings file with the extension .dvset. This contains your subject and client lists and various other settings. And don’t forget your dongle, or if you use an electronic licence key, make sure that the key will apply to the other computer.

12. How to find out more:
For more detailed information it is worth looking at the DVX2 User Guide for DVX2 Professional or DVX2 Workgroup. The link is at the bottom of the page, and the user guides are PDF files with over 600 pages. On the website http://www.atril.com there are also links to various videos, webinars and training courses, and also to the mailing list dejavu-l (under Support>Technical forum).

I already mentioned Steven Marzuola’s article on terminology databases. It is also worth looking at Nelson Laterman’s collection of tips and tricks for DVX1 (and even its predecessor DV3).
I am sure there are plenty of tips and questions which I have not covered, so I am looking forward to reading comments by my readers.

Wednesday, 30 November 2011

Kindle eReader: tool or toy?

Curiosity finally got the better of me, and I am now the owner of an Amazon Kindle eReader - the version with a keyboard, Wi-fi and 3G Internet access.

First impressions

I don't want to go through all the features - there are plenty of technical websites that do that (including Amazon's own website). I will focus on two aspects. Firstly, what have I noticed about its usability in practice over the first few days? And secondly, how useful will it be for me as a translator?

Let's start with a couple of negative points. Although the text is crisply defined and can quickly be adjusted to different sizes, the background is rather grey. I knew in advance that the Kindle screen is not "backlit", so it needs daylight or artificial light to read. But comparisons with the legibility of text on paper are only partly true, because the background is darker than paper. Reading it in a dimly lit room is rather difficult, so you need a reading light or a clip-on battery light. With the right lighting, however, it is easy and pleasant to read.

Turning the pages of a book is easy and quick, especially compared with printed books. In other respects, however, navigation is slightly clunky and takes some getting used to. There is no mouse or touchpad, and this Kindle doesn't have a touchscreen. To move around on the page, there are 4 tiny little arrow keys, and to move to a word in the middle of the page you have to press the down and left/right keys several times. I suppose I am spoiled by my other equipment: desktop PC with a mouse, laptop/netbook with a trackpad or mouse, smartphone with a touchscreen. So my first impression of the Kindle keyboard is rather like time travel - as if I were moving back to a slightly older technology.

Some of the ebooks that I have downloaded are even more difficult to navigate. One of the things I want to do with the Kindle is to read the Bible. I have checked a number of Bibles in both English and German, and incredibly I find that many of them have no table of contents at all. The Bible is not the sort of book that you read sequentially from front to back, so a table of contents is essential. I have found one or two that I can use, but the selection of properly indexed Bibles is very small indeed.

One feature of this Kindle is the free Internet access over the 3G network in all of the countries that I am likely to travel to. This feature is mainly designed to let me access the Amazon store when I am on the road, but the Kindle also has a rudimentary browser (which Amazon calls "experimental"). I have tried it, and I am really able to access my own e-mail account with this browser. But operating a browser with only arrow keys and no mouse feels rather clumsy. It is easier, faster and more pleasant to check e-mails and the Internet with my smartphone, in spite of the smaller screen. So I will hardly use the "experimental" browser in Germany, where I have an Internet flatrate on the smartphone. But it will be useful, for example, when I visit the UK and am not within reach of a Wi-fi access point.

Kindle for translators?

On my Kindle I have three free monolingual dictionaries (Oxford Dictionary of English, New Oxford American Dictionary and Duden Universalwörterbuch). Here, the indexing is excellent. I can choose one of them as my default dictionary, and when I am reading on the Kindle I can look words up directly from the text. Or I can open one of them from the menu and search in the dictionary, and even turn the pages to check out entries before and after the keyword I have entered. A couple of times during the last few days I have used these dictionaries to check terms in both German and English in the course of my work. I will probably also download a thesaurus for English, and one for German, too.

Amazon's Kindle shop offers various bilingual dictionaries, although most of them seem to be targeted at general users rather than professional translators. There may be some specialist dictionaries worth buying - for example I am currently checking the free sample of an illustrated bilingual engineering dictionary. The Kindle Shop could also be a useful source of monolingual specialist literature. There are dictionaries in either language for subjects such as law, property/construction and many others. It also offers the text of German laws for a very moderate price.

Another feature of the Kindle is that I can send my own documents to it in various file formats. This could be useful for anything I need to refer to during my work (source documents, abbreviation lists, background texts etc.). To test this function, I sent the DVX2 manual to my Kindle. It is a PDF file which is over 600 pages long, and the table of contents is not indexed for the Kindle, so navigation is limited. But I entered the search term "DeepMiner", and it jumped through the manual from one instance to another until it found the section that actually explains how this function works. I was then able to rotate the screen to wide format and adjust the size so that I could read it reasonably well. The display is not in colour, and navigation is more clumsy than on a desktop or laptop computer, but for some purposes this function could be useful.

The classic use for the Kindle, of course, is to read books from start to finish. This works well, and it is convenient to have a selection of books in just one relatively lightweight device which claims to be able to store 3,000 books or more (especially when travelling). Only time will tell whether I use my Kindle mainly for leisure reading purposes, or whether it really becomes a regular part of my workflow.

Friday, 11 November 2011

DVX2 screenshot gallery

At first sight, the screen of the Translation Memory program DéjàVuX2 (DVX2) is just a mass of boxes, a chaotic pattern of vertical and horizontal lines. What are they all for? Where in this enormous jigsaw puzzle can I find the text I want to translate? What other information is provided on the screen, and how is it helpful? The best way to explore this is with screenshots.

The classic layout
When you start working on a project with DVX2, the screen will probably look something like this. The pane at the top left is the working area. The left column is headed "German" - that is my source language. The right column, English (United Kingdom), is where my translation goes.

At the bottom left and bottom right of the screen I can see my reference material. At the bottom right I have terminology suggestions ("AutoSearch Portions"), and at the bottom left I have similar sentences ("AutoSearch Segments"). The top right ("Project Explorer") shows me the files in the project. When I am working on the translation, I normally hide this pane so that I have the full window height for the terminology.
There are various ways to personalise this layout. I can change the font and type size in the various windows, and I can also change the arrangement of the different panes in the working window.
My personal layout
Modern monitors, laptops and netbooks tend to have a wide screen. There is not much space to display elements above each other, so it is sometimes better to display the elements side by side. Therefore, my normal DVX2 screen looks like this:
In this "tramline" layout, the working area is in the middle of the screen and the reference material is arranged to the right and left. It provides more context (i.e. the text before and after the active sentence). The shorter lines could be a disadvantage for longer sentences, and especially on smaller screens. The above screenshot is taken from my 22" monitor. On my 10" netbook, this layout is rather more cramped, although it would be just about workable:
One way to make the lines longer in the working area is to work in a separate text area at the bottom of the screen and to split this text area vertically (Tools>Options>Environment). The active sentence is highlighted in the grid, but the working area is now at the bottom, i.e.:
I often get jobs with very long sentences, and sometimes the reference pane on the left is empty for most segments. In such jobs, I can simply hide this column, which gives me longer text lines even without using the separate text area:
Hide and display
In the last screenshot, note the little tabs on the left and right of the screen. They are "mouse-over" tabs. If I want to have a quick look at "AutoSearch Segments", I simply move the mouse over the tab, and the AS Segments pane opens up, but closes again when I return the mouse to the main grid.
Note also the little drawing pin icon at the top right of the "AS Portions" pane. This is a three-way switch for the display of this pane. It can either be fully displayed, as it is here, folded away like the "AS Segments" pane, or it can hover as in the mouse-over function. The combination of the tabs and the drawing pin icons takes a bit of practice, but it helps me to be flexible in using the screen layout.
Smaller details
There are a number of smaller details in the screen layout which can be useful.
The top of the DVX2 window shows the name and path of the current project. For example, the project I used for these screenshots is on drive D at the location shown.
These six icons are in the middle of the bottom edge of the DVX window. Mousing over them displays what they mean - here I had the mouse over the first icon (AutoWrite). The background colour shows me whether the function is on or off. Here, for example, AutoWrite, AutoAssemble, AutoPropagate and AutoCheck are enabled, but AutoSearch and AutoSend are disabled. These functions can also be switched on or off via Tools>Options>Environment, but the icons are quicker.

This is the area above the working part of the grid, and it contains a few hidden details. The grid language heading boxes (here "German" and "English") switch between alphabetical and chronological view of the project sentences. The language field with the flag has a little arrow to the right, which leads to a list of the target languages in the project (useful for project managers, but not usually for freelancers like me). The box "All segments" also has a little arrow, which opens up a list of types of sentence (all fuzzy matches, all exact matches etc.). The empty box on the left is a row finder. If I know the number of a segment, I can type it here, and DVX2 jumps to that segment (useful if I am proofreading and notice that a segment needs more work when I have finished proofing - I simply jot down the number and jump to the segment afterwards).
The tabs above this row show the name of the files which I have opened, so I can move to another file simply by clicking the tab. That in itself does not sound special. But these tabs can also be used to display files side by side (or one above the other). I can then compare my work on two files in context, for example like this:
This article only looks at the main grid, in other words the screen which I usually see when I work on a project. It does not explore the menu or any of the subsidiary screens, nor does it examine the efficiency of the many functions of the program. But I hope that this visual summary gives a general impression of the working environment.

Wednesday, 19 October 2011

Deep mining with Déjà Vu X2

"Déjà Vu" is a translation memory program created by the company Atril. It stores my previous translation work in databases and uses these databases to help me in every new translation job. There are three types of database. The "translation memory" (TM) contains the sentences in my source language together with my translations of these sentences. My main TM has about 385,000 sentence pairs in my two languages (German and English), so it is effectively an archive of all the work I have done since I started using Déjà Vu in late 1999. The "termbase" (TB) has terminology items which I have entered. My main TB has about 54,500 terminology pairs. In addition, there is a "lexicon" for each project, which is a place to put proper nouns, client-specific terminology etc. When I work on a new translation project, the program calls on these databases to offer as much help as possible. Sometimes this enables me to work much faster on a project. But usually I have projects with long and complicated sentences, especially contracts, so the speed gains are usually more modest. The main benefit of Déjà Vu for me is as a tool for quality which enables me to be more consistent in my work.

Over the last 12 years I have seen three generations of the program. The first version was known by the abbreviation "DV3". The next generation, DVX, was released in May 2003. The latest version is DVX2, which was released in May 2011.

Each new version has new features. A list of new features in DVX2 can be found here. One new feature which has puzzled many people is "DeepMiner". The theory is that it uses both the TM and the terminology databases to retrieve even more material. But how does it work in practice? There is a training video which uses an extremely simple example to show cross-analysis between the sentences "I have a brown dog" and "I have a black dog" when translating them into French.

So far so good. In practice, however, my sentences are never as simple as this example, and the size of my databases means that DeepMiner has to work much harder. As a result, using DeepMiner on a largish project with big databases can be very slow. And in my experience, DeepMiner is sometimes not helpful because it tries to be too clever and reconstruct the solution from similar sentences in the TM, and in the process it may overlook what I have in my termbase and lexicon. Thankfully, it is easy to switch the DeepMiner function on or off.

So how helpful is this new function? To illustrate this, let's look at one example sentence from a complicated German land purchase and partitioning contract in two alternative versions: with and without DeepMiner:

My translation:
I make the following declarations not in my own name, but as a manager with power of sole representation of ...

Looking at the first half of the sentence, where do the phrases "my own name" and "the following declarations" come from in the example with DeepMiner? They are not in the terminology hits for this segment, and there is no whole sentence match. But the TM has many matches containing "die nachstehenden Erklärungen" and the translation "the following declarations" (although "nachstehend" on its own is only in the TB as "hereinafter"). The first three words "my own name" seem strange at first sight. Somehow, DeepMiner seems to have found a correlation between the words "ich ... im eigenen Namen" and the English "my own name", in spite of the fact that the TB entries which use "im eigenen Namen" only offer the English "its own name" and "his own name".

At least in this example, DeepMiner offers solutions which go beyond the conventional assembly and pretranslation routines in the previous version of DVX. In my experience, it is still a matter of trial and error - sometimes it finds surprisingly good suggestions, but sometimes it is not really helpful. One possible workflow to get the best of both worlds is to "Pretranslate" the whole file with DeepMiner activated and then, if the solution is not helpful, to "Assemble" the individual sentence without DeepMiner. To do this, the settings for Pretranslate are:

And the settings for Assemble (under Tools>Options>General) are:

I am still experimenting to find out how DeepMiner can be used to best advantage, so perhaps I will be able to add more insights at a later date. Before too long (hopefully) I will comment on some of the other features of DVX2 such as AutoWrite, the information design options in the variable grid layout etc.