Jump to content

Help talk:CirrusSearch/2015

Add topic
From mediawiki.org
Latest comment: 10 years ago by Elvey in topic Search text not found


How does the search API for MediaWiki work?

[edit]

I am using media wiki search API i.e. http://en.wikipedia.org/w/api.php?action=query&list=search&format=json&srsearch=Taj+Mahal+Agrain one of my project.

When I am trying to search for Taj Mahal Agra, my search results does not show any result for Taj Mahal although its one of the very popular place of India.

Infact there are no actual search results are returned, instead some suggested results are returned and it is not even possible to identify that the returned results are not the actual results and they are the suggested one instead.

Expected was, that the API to return actual search result with atleast for Taj Mahal as the first element in the resulting response.

And my concern is that I need to search with the string "Taj Mahal Agra" and not with only "Taj Mahal".

Please advice.

Thank you in advance. Krishdamani (talk) 13:25, 7 January 2015 (UTC)Reply

When I run that query against the API I see the first 10 of 405 results for any pages containing the words "Taj", "Mahal", or "Agra" sorted by relevance. The API result matches the same query via the search box on the English Wikipedia for me. If you query instead for "Taj AND Mahal AND Agra" you will get only pages that contain all three words and en:Taj Mahal becomes the first result. The default "or" behavior and the resulting scoring may be what is confusing you in this case. BDavis (WMF) (talk) 02:25, 17 February 2015 (UTC)Reply
Should we really expect users of the encyclopedia to expect "or" as the default? I think that is a very bad assumption. In case we continue to make it, at least there should be some sort of notification before or after searching. Kdammers (talk) 08:35, 21 February 2015 (UTC)Reply
Using a default "OR" conjunction is fairly common to the point of being nearly universal in web based consumer facing search. This choice is typically based on the assumptions that the scoring model will surface better matches near the head of the result list and that showing more results is preferred to showing fewer. A Google search for example follows the same behavior.
Google however also has a very large corpus of hinting that helps them understand that "Taj Mahal" is a phrase even though it was entered without explicit phrase quoting. You can see the results of this type of change on CirrusSearch from the default 'Taj OR Mahal OR Agra' parsing by entering instead '"Taj Mahal" Agra'. This change also brings
en:Taj Mahal
to the top of the results list in the same way that using and explicit AND conjunction does. Interestingly, in this case
en:Taj Mahal
is the first result if Agra is dropped from the query entirely. This is likely due to the increased sorting boost given to title matches. BDavis (WMF) (talk) 19:11, 22 February 2015 (UTC)Reply

Wiki markup in search results

[edit]

Is there a way to get formatted text in search results instead of wiki markup? 140.108.1.12 02:47, 16 January 2015 (UTC)Reply

You can search the wikitext with "insource", as noted in the help page.
Or do you mean in the search snippets shown on Special:Search itself? Nemo 12:13, 16 January 2015 (UTC)Reply

Jargon

[edit]

Can we avoid the jargon? For example, I have no idea what "null edit" means. Write simply. Yours, GeorgeLouis (talk) 19:38, 22 January 2015 (UTC)Reply

DO NOT USE

[edit]

This feature apparently needs a rewrite from scratch with some simple syntax tricks explained on the search page itself. Notably about 99% of my searches are about meta-stuff (here covered by an external Google search, folks obviously know what's good and what's hopeless.) Otherwise I want the common features like "inurl:" (instead of full text), file types for files, "*" anywhere for "whatever", quotes for literal, + for required, the works.

Background: I have a garbled filename "Catedral_de_Skálholt,_Suðurland,_Islandia,_2014-08-16,_DD_142.JPG" on my disk (downloaded from commons for a review), and now I tried to find the corresponding "File:Catedral_de_Sk*.jpg" again, miserable FAIL for the abomination claiming to be a search engine. Be..anyone (talk) 01:08, 2 February 2015 (UTC)Reply

Works for me Nemo 06:38, 2 February 2015 (UTC)Reply
Works also for me, good trick. Be..anyone (talk) 13:50, 2 February 2015 (UTC)Reply

Brackets in search suggestions

[edit]

The old search seemed to ignore the absence of brackets in text typed into the search box, and gave search suggestions for articles that included brackets in their titles. This doesn't seem to work any more. For example, I used to be able to type in "sweet band" and "Sweet (band)" appeared as a suggestion. Now the article by that name only appears in search suggestions if I type in "sweet (band)". Viennese Waltz (talk) 15:18, 2 February 2015 (UTC)Reply

If you don't know whether the title contains parentheses you can also enter "sweet band ~", which will include that article in the suggestions (albeit a bit down the list). SebastianHelm (talk) 17:22, 4 February 2015 (UTC)Reply
Thanks for the response, but sorry, that doesn't work for me. To be clear, I'm not referring to the list of search results running down the main page. I'm referring to the suggested hits that appear in an auto-complete fashion when I start typing in the search box. When I type "sweet ba" I get a bunch of suggestions, but when I go on to type "sweet ban" the suggestions disappear, and typing ~ does not bring them back. In any case, although I'm grateful for your input, I'm not looking for an alternative. I'm asking why the previous functionality doesn't work any more, and whether it can be brought back. Thanks again. Viennese Waltz (talk) 08:30, 5 February 2015 (UTC)Reply
Hey! Its not supported because its not a feature I realized existed. Its a good, obvious feature. Its so obvious that I used it in the old search without realizing it. Anyway, I've filed it as T89201. It should be reasonably simple to implement when we next get a chance to work on Cirrus again. NEverett (WMF) (talk) 02:01, 11 February 2015 (UTC)Reply
Wonderful, thanks very much. Viennese Waltz (talk) 10:37, 12 February 2015 (UTC)Reply

exclude redirections ?

[edit]

Is there a way to exclude redirection pages from search results ? For example, I'd like to isolate only real articles in this search. Thanks for your help. Tomates Mozzarella (talk) 21:31, 5 February 2015 (UTC)Reply

No one to help me ? ;-( Tomates Mozzarella (talk) 14:49, 20 February 2015 (UTC)Reply
You can user -insource:/REDIRECT/ and similar, of course, but it's not particularly clean. Nemo 07:48, 21 February 2015 (UTC)Reply
I've filed the feature request at phab:T90807 ("Option to exclude redirection pages from search results") Quiddity (WMF) (talk) 23:02, 25 February 2015 (UTC)Reply

How to search for "and" (logical conjunction) of multiple terms?

[edit]

The documentation is not clear on how to search for logical conjunction of terms ("a" and "b"). In particular, I would like to search for instances of a particular template that contain a certain term. E.g.: instances of "hastemplate:cite conference" that contain "chapter=". Any suggestions? 50.251.218.25 22:37, 5 February 2015 (UTC)Reply

I think insource will work for this? Try hastemplate:"cite conference" insource:"/chapter=/"
https://en.wikipedia.org/w/index.php?title=Special%3ASearch&profile=default&search=hastemplate%3A%22cite+conference%22+insource%3A%22%2Fchapter%3D%2F%22&fulltext=Search Quiddity (WMF) (talk) 07:11, 7 February 2015 (UTC)Reply

Obfuscated Wikipedia sic doesn't work as I expect

[edit]

I don't know if this is a bug or a feature, or something that needs fixing in WP template sic, but obfuscated sic in Wikipedia (WP) doesn't work as I'd expect. {{Sic|super|cede|hide=y}} displays in a WP article as "supercede" but the standard WP search doesn't find it, thus preventing it being found and blindly "corrected". But the new search does find these hidden instances. For example, a search for "supercede" with the old and the new search on Wikipedia includes obfuscated results for the new, but not the old, search. 17:09, 6 Feb 2015, Wikipedia editor pol098. Wikipedia editor pol098 17:17, 6 February 2015 (UTC)

It’s a feature. See #What's improved?, third point. Tacsipacsi (talk) 18:57, 6 February 2015 (UTC)Reply
Thanks for reply. I think this is "Expanding templates, meaning that all content in an article that's in a template is now reflected in search results." WP Pol098 20:24, 6 February 2015 (UTC)

search times

[edit]

Why are my searches taking so much more time on the new search engine now than they were on the older search engine? Any help on this will be greatly appreciated. Roverfan77 (talk) 09:37, 7 February 2015 (UTC)Reply

Sometimes the new search engine allows more complex queries. Can you make examples of slow queries? Nemo 21:17, 8 February 2015 (UTC)Reply

Case sensitivity

[edit]

Is CirrusSearch really unable to do case sensitive searches? This seems like a basic requirement for any "advanced" search, and it has been implicitly raised by several other users here, most notably by Mikhail Ryazanov in #"Really" exact matches. Is there a bugzilla case or phabricator task for this? SebastianHelm (talk) 17:39, 9 February 2015 (UTC)Reply

For case sensitivity you can use the regex "insource" search. Nemo 19:03, 9 February 2015 (UTC)Reply
Thanks, that resolves my question. When I just read the description at Help:CirrusSearch#insource:, I noticed that "The version with the extra i runs the expression case insensitive and is even less efficient." How about adding a remark there that recommends using just simple /.../ instead? SebastianHelm (talk) 01:26, 10 February 2015 (UTC)Reply
I don't understand your suggestion. Regex search is slower, so it should only be used with a reason. Regexes are case sensitive. Rarely, one can need a case insensitive regex, which is notoriously slow. If you need a simple case insensitive search for a word, then better not use the regex search. Nemo 09:37, 10 February 2015 (UTC)Reply

Edit date/time

[edit]

Is it possible to search for text by the date/time it was added? SebastianHelm (talk) 18:36, 9 February 2015 (UTC)Reply

That's not possible with the (old or new) site search, as it only searches the current revisions.
There is an external tool which aims to support this, although it is fairly complicated (for me at least) to use initially (I forget how it works, and have to re-learn via experimentation, each time I use it). See en:WP:WikiBlame for links and details. Quiddity (WMF) (talk) 01:17, 13 February 2015 (UTC)Reply

search for mobile MediaWiki the same time

[edit]

Search Cirrus just opens two more useless search fields but no results on the otherwise blank page. 166.216.194.e89 05:04, 15 February 2015 (UTC)Reply

What is "mobile MediaWiki"? Can you post the URL and if possible a screenshot of what yoou are seeing? Nemo 14:06, 15 February 2015 (UTC)Reply

Small fix suggestion for clarity

[edit]

Searching for: Female writers from the Netherlands, yielded: category:Female writers from the Netherlands. This confused me, and I would rather see: Female writers from the Netherlands.

If it wasn't for the box that said: This is a new search engine, I would have gone away. Wouter Drucker (talk) 07:33, 24 February 2015 (UTC)Reply

Hi, thanks for the comment. You are talking of commons:Category:Female writers from the Netherlands, right? commons:Female writers from the Netherlands doesn't exist, can you clarify why you'd prefer to see a title which doesn't exist? Nemo 12:03, 26 February 2015 (UTC)Reply

Feature suggestion

[edit]

On the phabricator pages folks discuss some obscure feature related to file uploads on phab. I vaguely recall that I added links to two images on phab as "other versions" on a commons file. So where was this, how can I find it again? Maybe Special:Contributions should offer a search limited to all pages edited by the given user. Be..anyone (talk) 07:53, 24 February 2015 (UTC)Reply

I don't know if worth it, but this could be feasible, "simply" dumping the history into ElasticSearch. Even just usernames would end up being huge, though. Nemo 06:47, 25 February 2015 (UTC)Reply

Phrase Search Specifics

[edit]

I have been using Wikipedia Search as the basis of my linguistics research: entering a series of phrase searches incorporating a specific noun (e.g. 'boy') and then counting the results. These are the actual search patterns: "of the boy"; "the boy's"; "of a boy"; "a boy's"; "of boy"; "boy's"; "boy". And the same examples, but with a wildcard inserted before 'boy' in each phrase. The specific requirements are: - exact phrase only - recognition of the apostrophe - wildcard representing any (one) word

Will the move to Cirrus Search affect this, and if so how will I need to adapt my methods? Advice and guidance much appreciated. Kevin KevinGlover (talk) 11:11, 24 February 2015 (UTC)Reply

If you want to match the apostrophe, you must use regexp, which is slow. FriedhelmW (talk) 16:51, 24 February 2015 (UTC)Reply

hassle,annoying

[edit]

Ok..I just want to say, since the change I have to make my searches very generalized,,Then search and hunt all over..Don't you think it could be more exact the first time????? Irritated in Missouri and is looking for a better site for information. Kittyennis (talk) 18:40, 27 February 2015 (UTC)Reply

Thanks for the comment, what do you mean by "very generalized"? That specific keywords don't work? Nemo 08:57, 28 February 2015 (UTC)Reply

CirrusSearch installation

[edit]

Hello! I am sorry for my noobery on the subject...

I wanted to do a specific search "purple +incategory:"Public domain"

and it said I there was a new search engine. "CirrusSearch"

I did a bit of research to see that I needed to download this folder into my (i am on mac) "/library/web server/documents" folder

So i did that.

Then I restarted the browser and tried the search and it seems like it still isn't working as expected.

Am I supposed to so something with the files once they are extracted into the folder?

I really am not a developer, so any help would be so appreciated! Calvin200001 (talk) 11:27, 8 March 2015 (UTC)Reply

Please read Extension:CirrusSearch, sections 'Dependencies', 'Installation' and 'Configuration'. FriedhelmW (talk) 16:48, 8 March 2015 (UTC)Reply
FriedhelmW (talk) 16:48, 8 March 2015 (UTC)Reply

Search for this but not that

[edit]

On the previous search you could search for "forth" - "bridge" and that would list articles that contained the word forth but not the word bridge. We seem to have lost that functionality, could we add that functionality to the new search or revert to the old one please? WereSpielChequers (talk) 19:17, 12 March 2015 (UTC)Reply

It works. You shouldn’t add a space between the “-” and the word: forth -bridge. Tacsipacsi (talk) 19:34, 12 March 2015 (UTC)Reply
thanks, now how do we get that into the documentation? WereSpielChequers (talk) 20:34, 13 March 2015 (UTC)Reply
What’s the difference? This is in the first paragraph: “This page describes the features that are new or different compared to the past solutions.” Tacsipacsi (talk) 21:11, 13 March 2015 (UTC)Reply
The old document is neither linked to nor available. Cpiral (talk) 22:24, 26 July 2015 (UTC)Reply

Incomplete search results

[edit]

When performing this search: https://commons.wikimedia.org/w/index.php?title=Special%3ASearch&profile=images&search=insource%3A%2F\{\{Information.*\{\{Information%2Fi&fulltext=advance i am expecting that commons:file:YAMAHA (headquarters 1).jpg is included as well. Why it is not the case? Thanks in advance. Aschroet (talk) 04:58, 29 March 2015 (UTC)Reply

It’s case sensitive by default and the Japanese template name is like information. Search for insource:/…/i (see #insource: section). Tacsipacsi (talk) 07:19, 29 March 2015 (UTC)Reply
The trailing "i" is already there. Even if i modify the query it does not work: https://commons.wikimedia.org/w/index.php?title=Special%3ASearch&profile=images&search=insource%3A%2F\{\{.nformation.*\{\{.nformation%2F&fulltext=Search Aschroet (talk) 08:11, 29 March 2015 (UTC)Reply
Then I have no idea what’s wrong. Tacsipacsi (talk) 08:21, 29 March 2015 (UTC)Reply
Neither do I. I've filed this as T94830 and placed it near the top of the backlog. NEverett (WMF) (talk) 13:16, 2 April 2015 (UTC)Reply
thanks. Aschroet (talk) 13:57, 2 April 2015 (UTC)Reply
Beside the mentioned false negative file, there is also a false positive file in the results, namely File:Anemochory.jpg resp. File:Vincetoxicum rossicum seeds.JPG which does not contain two {{Information. Aschroet (talk) 15:14, 29 March 2015 (UTC)Reply
I don't get those in the search results. NEverett (WMF) (talk) 13:10, 2 April 2015 (UTC)Reply
i can confirm that this is no longer the case but it was when i was running this query that day. I would ignore this and focus on completeness. Aschroet (talk) 13:57, 2 April 2015 (UTC)Reply

Impact of word order in two-words search query

[edit]

m:Special:Search/engineering reorganization has m:Wikimedia Foundation Engineering reorganization FAQ/en as 1st result, while m:Special:Search/reorganization engineering has it in 24th position, preceded by a bunch of search results which only have "reorganization" in the snippet, while "engineering" doesn't appear at all or is just in a section: see [1].

I gather that word proximity in the document doesn't matter if the two words were searched in the opposite order as they appear, right? Is this expected? Can it be changed? Nemo 17:01, 28 April 2015 (UTC)Reply

The proximity in the document doesn't matter unless you put the query words in quotes. Then you can add a proximity variable to that single filter. It works backwards and forwards, but doesn't make the last word proximate to the first. Two search terms with no quotes is two filters, and a bunch of page-ranking rules. Cpiral (talk) 20:16, 23 July 2015 (UTC)Reply

Searching for a bolded left parentheses?

[edit]

Any ideas on searching for a Left Parentheses which is bolded when displayed?

So '''abc''' '''(def''' would qualify, but abc''' '''(def''' would not. Naraht (talk) 06:34, 8 June 2015 (UTC)Reply

It's impossible. Regex engines are single pass, and so cannot behave earlier on what is counted later. Cpiral (talk) 20:11, 23 July 2015 (UTC)Reply
But a Regex engine could keep track of whether at a specific point in the pass it is in a bolded state or not, right? Naraht (talk) 13:22, 2 September 2015 (UTC)Reply
Nope. That state depends on the future of its single pass, which is uncertain to contain the closing marks. I supposes a look-ahead assertion would work, but we don't have that option here.
The best we can do is find three tick marks, any number of characters that are not a single tick mark, followed by the left parenthesis.
insource:/'''[^']*(/
The look ahead assertion of some regex search engines can expand that from a single tick-mark character, but even so, it fails your criteria, as it only compares two at a time, not three. You'd need to process the text one line at a time. But we can't do that with Cirrus Search.
We have no metacharacter for newline. We do have the dot . for newline, but that matches any character. We also have regex like the above [^'] but it too matches the newline plus all characters but that tick.
Now, the three ticks will bold everything up to the end of the line, even without its closing three tick marks, and since we have no newline metacharacter, we can't discern this case:
'''bold text
( not bold ( ''' bold again

bold text ( not bold ( bold again

Cpiral (talk) 00:17, 3 September 2015 (UTC)Reply

Search does not fuzzily match

[edit]

Do I have a configuration problem or am I misunderstanding the software's intended behavior?

My wiki has many pages containing the word "install." When I search on "instal", I expect it to fuzzily match all the instances of the word "install," or to suggest a correction. However, the search simply returns empty. Markmichaelh (talk) 19:23, 11 June 2015 (UTC)Reply

You need to search for
instal* FriedhelmW (talk) 19:38, 11 June 2015 (UTC)Reply

insource: special characters like \s \w \n

[edit]

They are regexp standard, but not supported here, right? Are they going to be? Bultro (talk) 09:09, 30 June 2015 (UTC)Reply

The # sign

[edit]

The # sign is a metacharacter, as stated on the "syntax" link under "insource:", but it is not covered anywhere. How is the metacharacter # used in CirrusSearch? It's reserved? Cpiral (talk) 23:28, 9 July 2015 (UTC)Reply

To be more specific: at it says at https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-regexp-query.html#regexp-syntax that:
"The standard reserved characters are:
. ? + * | { } [ ] ( ) " \
"If you enable optional features (see below) then these characters may also be reserved:
# @ & < > ~
In the same reference it goes on to describe each optional feature except #. I MediaWiki-Elastica does have the optional features enabled because they all work as documented.
I was unable to find anything about ElasticSearch and the # sign on the Web. Experimenting, I could conclude that the # sign does not function as a comment in the regexp, and to be honest I can't think of any way the CirrusSearch regexp repertoire is lacking. I was just curious about what I may be missing. So although I don't see needing it as a metacharacter, I do see needing the documented reference to the # sign to be fulfilled somewhere. Cpiral (talk) 19:41, 11 July 2015 (UTC)Reply

CirrusSearch problem

[edit]

In the search box at Wikipedia this should work, but does not:

insource:/\{[Ii]nfobox unit.*inunits1 *= *(\{\{)?[^#<>[\] {}]+\|?[^|][0-9]/ prefix:A

But 1) It works if prefix is at least "prefix:As", (for "Astronomical unit")

And 2) It also works if hastemplate:"infobox unit" is added.

The two things that make it work don't necessarily make sense, do they? Cpiral (talk) 20:36, 14 July 2015 (UTC)Reply

T112726 WAS OPENED ON THIS. Cpiral (talk) 06:26, 24 September 2015 (UTC)Reply

CirrusSearch problem #2

[edit]

hastemplate:"Val" insource:/\{\{[Vv]al\|[^}]*m\// prefix: :

runs in a snap. NOW ADD AN "s" at the end for "m\/s", and it takes so long that it times out, saying, consistently

" There were no results matching the query."

That is a problem, because it is obviously in error: There are plenty of pages that match that search without the "s", many of them showing the "s" clearly in the match. It should report a timeout error or some other kind of error.

NOW ADD ONE MORE CHARACTER in the pattern, such that we no longer have the "/s at the end" problem:

hastemplate:"Val" insource:/\{\{[Vv]al\|[^}]*m\/s[|}]/ prefix: :

then it runs in a snap again.

Mainspace has 1275 pages with template Val.

Userspace has 626 pages. (Change to `prefix:User')

That can't possibly be too many pages to "grep" through.

Userspace has 10 pages that match the target: Val .*m/s

Mainspace has 73 pages that match the failed search. (It works by adding another character to the pattern.)

This seems to be a throttle limit, because it takes a long time (it's slow), but most telling, it works for prefix:User:. The only difference being the number of pages. It barely works for user space. But then why not an error msg?

I refer you to https://discuss.elastic.co/t/how-to-protect-an-es-cluster-from-searches-that-would-kill-it/25487 Cpiral (talk) 17:23, 16 July 2015 (UTC)Reply

T12725 was opened for this Cpiral (talk) 06:22, 24 September 2015 (UTC)Reply

find

[edit]

The following discussion is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.


The discussion above is closed. Please do not modify it. No further edits should be made to this discussion.

Prefix:-

[edit]

Prefix:- gives entries that start with - or ' . How do I just get those that start with -? Naraht (talk) 01:51, 23 July 2015 (UTC)Reply

It's impossible.
Among the 446 search results on WP were one -F and two 'F. So Prefix is matching the dash you gave it to the single quote. Prefix does a character-wise, literal-character search, so currently it's impossible.
It's also impossible to find titles that begin with double quotes. All the other non-alphanumeric characters produce right results with one exception: the underscore.
Prefix:_ (with an underscore) also gives entries that start with - or '. So currently prefix with either an underscore, a single quote, or a dash, all find the same 446 search results.
There's no way to filter out the unwanted, single-quote titles except to resort to offline processing of the search results it does give: Intitle does only word-wise searches, and insource// does not search the title index.
Important to note: I can ''create a page title that starts with a single or double quotes'' character in the title, but I ''cannot find that page with Prefix''. And so it is in error that Prefix ignores the single or double quote character. Cpiral (talk) 20:09, 23 July 2015 (UTC)Reply
1920 5centavos 2607:FB90:D63:1993:EC7F:2115:E44:CC8A (talk) 22:39, 7 August 2015 (UTC)Reply
Thanx for confirmation that it doesn't work. Naraht (talk) 13:24, 2 September 2015 (UTC)Reply
T112722 was opened for this Cpiral (talk) 06:18, 24 September 2015 (UTC)Reply

Operation timed out

[edit]

If I follow the Readme and try to initial update my index I get (after 1h) a timeout. But from my understanding, the php script already connects to elasticsearch to fetch the version ... it finds also the index (If I do start it twice and do not remove it from elasticsearch per hand).

No clue, why the index generation/check does not work .. but the rest seems to work

/var/www/html# php /var/www/html/extensions/CirrusSearch/maintenance/updateSearchIndexConfig.php

content index...

        Fetching Elasticsearch version...1.6.0...ok

        Scanning available plugins...none

        Infering index identifier...spicewiki_content_first

        Picking analyzer...english

        Creating index...

Unexpected Elasticsearch failure.

Http error communicating with Elasticsearch:  Operation timed out.

194.138.39.61 (talk) 13:33, 18 August 2015 (UTC)Reply

insource can't find repeating words

[edit]

The following discussion is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.


There is no way to filter insource:/"<big></big>"/. Filters for regex are very important to have.

The search "big big" will not look insource, but will find repeating words.

The search insource:"big" and insource:"big big big" are equivalent.

Unlike "big big big", which uses proximity zero, insource turns off proximity.

In effect insource can't find repeating words. That's lame.

Insource should at least set proximity to zero. Cpiral (talk) 00:29, 3 September 2015 (UTC)Reply

The discussion above is closed. Please do not modify it. No further edits should be made to this discussion.

morelike:

[edit]

The morelike: operator seems to run on its own, unlike how it was labeled/categorized by its former section heading "special prefixes", and put with namespace, which is a query prefix. (I just separated namespace away from morelike:, to its own section.)

Prefer-recent: and boost-template: seem to make no difference to order or quantity of results.

If there are no disagreements, I'd like to document that it is a sole query term; i.e. morelike: is not a filter or prefix or suffix or in between term; i.e. morelike can go with no other terms. Cpiral (talk) 23:12, 24 September 2015 (UTC)Reply

Translations issues on this help page

[edit]

Currently there is yet another translation issue on this Help page. Would somebody please fix it?

I propose we remove translations from this page for a while. Cpiral (talk) 01:27, 27 September 2015 (UTC)Reply

I've removed translations tags and markers, and even the <languages /> tags. Not sure about all those. (See T113907 for why.)
The intent is extensive editing. It needs it.
My criterion for removing translations were these priorities: 1) Editing ability 2) Searching ability 3)Translations. Cpiral (talk) 22:18, 27 September 2015 (UTC)Reply
This is ready for translation now, after 4 months? Please explain current status of this page. Kaganer (talk) 23:30, 29 January 2016 (UTC)Reply
This page reached a milestone recently -- "fully informed" -- and so I'm glad you asked. This page is also too rough. Very soon I will smooth out the accessibility, readability, usability problems, and then call for translations. Cpiral (talk) 08:20, 30 January 2016 (UTC)Reply
Nemo, by reverting all my edits and telling me to add them while translations tags are on the page, you seem to have overlooked MediaWiki's first general principle of translations "Avoid changes". Cpiral (talk) 00:29, 15 February 2016 (UTC)Reply
Just for the record, here is what has happened.
  1. Translate administration translated the page, in my opinion, prematurely.
  2. No translations admin responded to direct appeals and clearly stated intentions on the talk page.
  3. Consequently, a significantly updated version was contributed over five months that one translation administrator @Nemo bis has only just now got around to reverting for the sake of translations alone (not because of its content) and another translation administrator @Shirayuki has (understandably) denied to re-markup for translations.
Assuming we are all just trying to do our job, I would just like to point out two principles at stake here: 1) the spirit of collaboration 2) providing for and enabling "customers" with documentation. Cpiral (talk) 02:15, 18 February 2016 (UTC)Reply

problems to reindex by elasticsearch

[edit]

current my it-admin moved the virtuell machine with our mediawiki runs (win7 64bit).

first the elastic search process Needs 100% cpu - today the cpu is better and in an other log-file i found the Information the index will be demage.

so i open cmd-window with admin-rights and send following command:

C:\PHP> C:\PHP>php c:/mediawiki/extensions/CirrusSearch/maintenance/updateSearchIndexCon fig.php

and this is the report....

´╗┐content index...
        Fetching Elasticsearch version...1.4.4...ok
        Scanning available plugins...none
        Infering index identifier...eblwiki-eblw__content_first
        Picking analyzer...german
        Index exists so validating...
                Validating number of shards...ok
                Validating replica range...ok
                Validating shard allocation settings...done
                Validating max shards per node...ok
        Validating analyzers...ok
        Validating mappings...
                Validating mapping...ok
        Validating cache warmers...
        Validating aliases...
                Validating content alias...ok
                Validating all alias...ok
                Updating tracking indexes...done
general index...
        Fetching Elasticsearch version...1.4.4...ok
        Scanning available plugins...none
        Infering index identifier...eblwiki-eblw__general_first
        Picking analyzer...german
        Index exists so validating...
                Validating number of shards...ok
                Validating replica range...ok
                Validating shard allocation settings...done
                Validating max shards per node...ok
        Validating analyzers...ok
        Validating mappings...
                Validating mapping...ok
        Validating cache warmers...
        Validating aliases...
                Validating general alias...ok
                Validating all alias...ok
                Updating tracking indexes...done
                Deleting namespaces...done
                Indexing namespaces...
Unexpected Elasticsearch failure.
Elasticsearch failed in an unexpected way.  This is always a bug in CirrusSearch
.
Error type: Elastica\Exception\Bulk\ResponseException
Message: Error in one or more bulk request actions:

index: /eblwiki-eblw__general_first/namespace/-2 caused UnavailableShardsException[[eblwiki-eblw__general_first][0] Primary shard is not active or isn't assigne
d is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@2f8d8f73]
index: /eblwiki-eblw__general_first/namespace/3 caused UnavailableShardsException[[eblwiki-eblw__general_first][0] Primary shard is not active or isn't assigned
 is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@2f8d8f73]
index: /eblwiki-eblw__general_first/namespace/4 caused UnavailableShardsException[[eblwiki-eblw__general_first][1] Primary shard is not active or isn't assigned
 is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@7a9b6074]
index: /eblwiki-eblw__general_first/namespace/7 caused UnavailableShardsException[[eblwiki-eblw__general_first][0] Primary shard is not active or isn't assigned
 is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@2f8d8f73]
index: /eblwiki-eblw__general_first/namespace/8 caused UnavailableShardsException[[eblwiki-eblw__general_first][1] Primary shard is not active or isn't assigned
 is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@7a9b6074]
index: /eblwiki-eblw__general_first/namespace/12 caused UnavailableShardsException[[eblwiki-eblw__general_first][0] Primary shard is not active or isn't assigne
d is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@2f8d8f73]
index: /eblwiki-eblw__general_first/namespace/13 caused UnavailableShardsException[[eblwiki-eblw__general_first][1] Primary shard is not active or isn't assigne
d is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@7a9b6074]
index: /eblwiki-eblw__general_first/namespace/102 caused UnavailableShardsException[[eblwiki-eblw__general_first][0] Primary shard is not active or isn't assign
ed is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@2f8d8f73]
index: /eblwiki-eblw__general_first/namespace/103 caused UnavailableShardsException[[eblwiki-eblw__general_first][1] Primary shard is not active or isn't assign
ed is a known node. Timeout: [1m], request: org.elasticsearch.action.bulk.BulkShardRequest@7a9b6074]

Trace:
# 0 C:\mediawiki\extensions\Elastica\Elastica\lib\Elastica\Bulk.php(345): Elastica\Bulk->_processResponse(Object(Elastica\Response))
# 1 C:\mediawiki\extensions\Elastica\Elastica\lib\Elastica\Client.php(284): Elastica\Bulk->send()
# 2 C:\mediawiki\extensions\Elastica\Elastica\lib\Elastica\Index.php(147): Elastica\Client->addDocuments(Array)
# 3 C:\mediawiki\extensions\Elastica\Elastica\lib\Elastica\Type.php(187): Elastica\Index->addDocuments(Array)
# 4 C:\mediawiki\extensions\CirrusSearch\maintenance\indexNamespaces.php(56): Elastica\Type->addDocuments(Array)
# 5 C:\mediawiki\extensions\CirrusSearch\maintenance\updateOneSearchIndexConfig.php(290): CirrusSearch\Maintenance\IndexNamespaces->execute()
# 6 C:\mediawiki\extensions\CirrusSearch\maintenance\updateOneSearchIndexConfig.php(213): CirrusSearch\Maintenance\UpdateOneSearchIndexConfig->indexNamespaces()
# 7 C:\mediawiki\extensions\CirrusSearch\maintenance\updateSearchIndexConfig.php(51): CirrusSearch\Maintenance\UpdateOneSearchIndexConfig->execute()
# 8 C:\mediawiki\maintenance\doMaintenance.php(97): CirrusSearch\Maintenance\UpdateSearchIndexConfig->execute()
# 9 C:\mediawiki\extensions\CirrusSearch\maintenance\updateSearchIndexConfig.php(58): require_once('C:\\mediawiki\\ma...')
# 10 {main}


very interesting is the statement "Elasticsearch failed in an unexpected way. This is always a bug in CirrusSearch"

is there someone to help me ???

here a part of current elasticsearch log:

[2015-10-09 11:14:10,356][WARN ][indices.cluster          ] [Living Hulk] [eblwiki-eblw__content_first][2] failed to start shard
org.elasticsearch.index.gateway.IndexShardGatewayRecoveryException: [eblwiki-eblw__content_first][2] failed to fetch index version after copying it over
	at org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:158)
	at org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:132)
	at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)
	at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)
	at java.lang.Thread.run(Unknown Source)
Caused by: org.apache.lucene.index.CorruptIndexException: [eblwiki-eblw__content_first][2] Preexisting corrupted index [corrupted_6KkIhZscQ92OMn2XKGuIYw] caused by: CorruptIndexException[codec header mismatch: actual header=0 vs expected header=1071082519 (resource: BufferedChecksumIndexInput(MMapIndexInput(path="C:\elasticsearch\data\elasticsearch\nodes\0\indices\eblwiki-eblw__content_first\2\index\_ky.fnm")))]
org.apache.lucene.index.CorruptIndexException: codec header mismatch: actual header=0 vs expected header=1071082519 (resource: BufferedChecksumIndexInput(MMapIndexInput(path="C:\elasticsearch\data\elasticsearch\nodes\0\indices\eblwiki-eblw__content_first\2\index\_ky.fnm")))
	at org.apache.lucene.codecs.CodecUtil.checkHeader(CodecUtil.java:136)
	at org.apache.lucene.codecs.lucene46.Lucene46FieldInfosReader.read(Lucene46FieldInfosReader.java:57)
	at org.apache.lucene.index.SegmentReader.readFieldInfos(SegmentReader.java:289)
	at org.apache.lucene.index.IndexWriter.getFieldNumberMap(IndexWriter.java:864)
	at org.apache.lucene.index.IndexWriter.<init>(IndexWriter.java:816)
	at org.elasticsearch.index.engine.internal.InternalEngine.createWriter(InternalEngine.java:1487)
	at org.elasticsearch.index.engine.internal.InternalEngine.start(InternalEngine.java:277)
	at org.elasticsearch.index.shard.service.InternalIndexShard.performRecoveryPrepareForTranslog(InternalIndexShard.java:732)
	at org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:231)
	at org.elasticsearch.index.gateway.IndexShardGatewayService$1.run(IndexShardGatewayService.java:132)
	at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)
	at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)
	at java.lang.Thread.run(Unknown Source)

	at org.elasticsearch.index.store.Store.failIfCorrupted(Store.java:480)
	at org.elasticsearch.index.store.Store.failIfCorrupted(Store.java:461)
	at org.elasticsearch.index.gateway.local.LocalIndexShardGateway.recover(LocalIndexShardGateway.java:120)
	... 4 more

BY 100% CPU !!!!!

regards Jan JanTappenbeck (talk) 09:10, 9 October 2015 (UTC)Reply

Have you read? https://github.com/elastic/elasticsearch/issues/4798
That looks like a problem in ES for me, not with CirrusSearch, but I have to say, that I'm not a pro with Elasticsearch/CirrusSearch :P Florianschmidtwelzow (talk) 09:15, 15 October 2015 (UTC)Reply

Boost-templates is effective?

[edit]

The following discussion is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.


Has anyone gotten any effect from using boost-templates parameter?

I've opened T115562, but maybe it should be closed. Cpiral (talk) 07:30, 15 October 2015 (UTC)Reply

Boost-templates works fine. Unlike hastemplate:arg, which assumes the Template namespace, boost-templates needs an explicit namespace:
boost-templates:Template:pagename. Cpiral (talk) 20:53, 16 October 2015 (UTC)Reply
The discussion above is closed. Please do not modify it. No further edits should be made to this discussion.

Search text not found

[edit]

I searched Wikipedia for "Key Personnel" in the Template space. There were 48 results, but not included among them was [[Template:Clearwater Threshers]], which has contained this exact string since it was created in 2009. Colonies Chris (talk) 14:10, 17 November 2015 (UTC)Reply

Since "key personnel" does show up on the page, but not in search results, it is either a bug, a common misunderstanding about how the collapse state relates to being indexed or not, or a code-base compromise of sorts.
Template:Clearwater Threshers uses | state = {{{state|autocollapse}}}. There are about 230 other navbox templates with "key personnel" who set collapse state like that, and none of them show up in the 45 or so searchable ones.
See what the 45 searchable ones use for their collapse state that do show up in search results. Most use {{{state<includeonly>|autocollapse</includeonly>}}} and a few use {{{state|collapsed}}}. It's not what I'd guess, since collapsed does not show the terms "key personnel" on the page, yet it is collapsed that seems to work for search.
Just like the HTML semantics determine the difference between a word search and an insource word search, so template semantics should determine the difference between the collapse state (invisible words) and not.
Its a good question for investigation. Cpiral (talk) 07:49, 19 November 2015 (UTC)Reply
I'm guessing this is an undocumented feature rather than a bug.
If I'm right, it should be documented, and IMO probably shouldn't be the default behavior of state/autocollapse. Elvey (talk) 19:39, 24 March 2016 (UTC)Reply

Help documentation

[edit]

Within a few days, I'll be adding information to the help page, and restructuring the sections.

Concerning other CirrusSearch documentation: besides phabricator, and wikipedia, and the sites already linked from the help page, are there any other resources that describe the behavior of the search box and its parameters that I should know about? Cpiral (talk) 23:34, 18 December 2015 (UTC)Reply

The wikitext is given in pithy chunks to help ease translations.
Translators: please copy this to a subpage and mark that up.
Please do not translate this page directly, as it is needing further editing. Cpiral (talk) 00:54, 9 January 2016 (UTC)Reply