wikiteam

Commit Graph

Author	SHA1	Message	Date
yzqzss	ebac66f557	Update dumpgenerator.py	1 year ago
yzqzss	0be46c7427	quote `title`	1 year ago
yzqzss	90a64c6a22	make `requests.session` to use `--retries` value (default=5)	1 year ago
Pokechu22	97146c6f01	Use the same requests session for getting the wiki engine and checking API/index	2 years ago
Pokechu22	6668999658	Update User-Agent to latest Firefox	2 years ago
Pokechu22	cad7260d7c	Fix crash when the image description is missing for an image containing non-ascii characters title is already unicode, so we shouldn't need to decode it (and don't in generateXMLDump).	2 years ago
Pokechu22	5b3fc4ac7b	Pass requests session to mwclient This means it uses our configured user-agent, as well as any cookies.	2 years ago
Pokechu22	1af69ca147	Skip empty revisions when using --xmlrevisions Before, the download would die, and need to be resumed from the start.	2 years ago
Pokechu22	4a2cbd4843	Use `session.get` instead of `requests.get` in `getXMLHeader` `session.get` uses our configured User-Agent, while `requests.get` uses the default one.	2 years ago
Pokechu22	9b2c6e40ae	Fix truncation when resuming There already was code that looks like it was supposed to truncate files, but it calculated the index wrong and didn't properly check all lines. It worked out, though, because it didn't actually call the truncate function. Now, truncation occurs to the last `</page>` tag. If the XML file ends with a `</page>` tag, then nothing gets truncated. The page is added after that; if nothing was truncated, this will result in the same page being listed twice (which already happened with the missing truncation), but if truncation did happen then the file should no longer be invalid.	2 years ago
Pokechu22	43945c467f	Work around unicode titles not working with resuming Before, you would get UnicodeWarning: Unicode unequal comparison failed to convert both arguments to Unicode - interpreting them as being unequal. The %s versus {} change was needed because otherwise you would get UnicodeEncodeError: 'ascii' codec can't encode characters in position 0-5: ordinal not in range(128). There is probably a better way of solving that, but this one does work.	2 years ago
nemobis	d7b6924845	Merge pull request #408 from shreyasminocha/fix-resume-images Fix image resuming	2 years ago
Tim Gates	ecbcc6118e	docs: Fix a few typos There are small typos in: - dumpgenerator.py - wikiteam/mediawiki.py Fixes: - Should read `inconsistencies` rather than `inconsistences`. - Should read `partially` rather than `partialy`.	3 years ago
Shreyas Minocha	e55de36cb7	Fix image resuming	3 years ago
Nicolas SAPA	b289f86243	Fix getPageTitlesScraper Using the API and the Special:Allpages scraper should result in the same number of titles. Fix the detection of the next subpages on Special:Allpages. Change the max depth to 100 and implement an anti loop (could fail on non-western wiki).	4 years ago
Nicolas SAPA	e4b43927b9	Fixup description grab in generateImageDump getXMLPage() yield on "</page>" so xmlfiledesc cannot contains "</mediawiki>". Change the search to "</page>" and inject "</mediawiki>" if it is missing to fixup the XML	4 years ago
Nicolas SAPA	eacaf08b2f	Try to fix a broken HTTP to HTTPS redirect in generateImageDump() Some wiki fail to do the HTTP to HTTPs redirect correctly so try it ourself.	4 years ago
Nicolas SAPA	7675b0d17c	Add exception handler for requests.exceptions.ReadTimeout in getXMLPageCore() Treat a ReadTimeout the same as a ConnectionError (log the error & retry)	4 years ago
Nicolas SAPA	4a5eef97da	Update the default user-agent A ModSecurity rule block the old UA so switch to the current Firefox 78 UA.	4 years ago
Rob Kam	e6f4674b42	fix typo	4 years ago
Federico Leva	abd908914f	Adapt to some more Wikia wikis edge cases * Make it easy to batch requests for some wikis where millions of titles are really just one-revision thread items and need to be gone through as fast as possible. * Status code error message.	4 years ago
Federico Leva	7de75012d1	Fix merge of the getXMLRevisions() loop	4 years ago
nemobis	8a2116699e	Merge branch 'master' into wikia	4 years ago
Federico Leva	7289225d2c	Directly catch exception for page missing in getXMLRevisions() The caller cannot catch the PME exception because it doesn't know about the title. Just log the error here.	4 years ago
nemobis	e136ee5536	Merge pull request #372 from nemobis/wikia Avoid launcher.py 7z failures	4 years ago
Federico Leva	8c6f05bb54	Consider status code before content in checkIndex() and checkalive.py Fixes https://github.com/WikiTeam/wikiteam/issues/369	4 years ago
Federico Leva	9ac1e6d0f1	Implement resume in --xmlrevisions (but not yet with list=allrevisions) Tested with a partial dumps over 100 MB: https://tinyvillage.fandom.com/api.php (grepped <title> to see the previously downloaded ones were kept and the new ones continued from expected; did not validate a final XML).	4 years ago
Federico Leva	a664b17a9c	Handle deleted contributor name in --xmlrevisions Avoids failure in https://deployment.wikimedia.beta.wmflabs.org/w/api.php for revision https://deployment.wikimedia.beta.wmflabs.org/?oldid=2349 .	4 years ago
Federico Leva	b162e7b14f	Reduce the API limit to 50 for arvlimit, gaplimit, ailimit Avoids to crash on errors or warnings which some wikis return for bigger requests, like https://www.openkm.com/wiki/api.php (MediaWiki 1.27.3).	4 years ago
Federico Leva	d543f7d4dd	Check the API URL against mwclient too, so it doesn't fail later Change the protocol from HTTP to HTTPS if needed. Fixes: http://nimiarkisto.fi/w/api.php	4 years ago
Federico Leva	d1619392f4	Force the lxml factory to pass around unicode strings Not necessarily the most compatible with downstream XML parsers, but at least should ensure that we manage to write the XML file. The encoding declared in the header is not necessarily the same we get from the API. See also: https://lxml.de/FAQ.html#why-can-t-lxml-parse-my-xml-from-unicode-strings https://lxml.de/3.7/parsing.html#serialising-to-unicode-strings Fixes https://github.com/WikiTeam/wikiteam/issues/363	4 years ago
Federico Leva	6dc86d1964	Actually use the next batch from prop=revisions in MediaWiki 1.19	4 years ago
Federico Leva	2ba69b3810	Indent the number of revisions more, consistent with page title style	4 years ago
Federico Leva	8fef62d46e	Implement continuation for --xmlrevisions with prop=revisions in MW 1.19	4 years ago
Federico Leva	8b58599645	Merge branch 'xmlrevisions' of github.com:nemobis/wikiteam into xmlrevisions	4 years ago
Federico Leva	17283113dd	Wikia: make getXMLHeader() check more lenient Otherwise we end up using Special:Export even though the export API would work perfectly well with --xmlrevisions. For some reason using the general requests session always got an empty response from the Wikia API. May also fix images on fandom.com: https://github.com/WikiTeam/wikiteam/issues/330	4 years ago
Federico Leva	2c21eadf7c	Wikia: make getXMLHeader() check more lenient, Otherwise we end up using Special:Export even though the export API would work perfectly well with --xmlrevisions. May also fix images on fandom.com: https://github.com/WikiTeam/wikiteam/issues/330	4 years ago
Federico Leva	131e19979c	Use mwclient generator for allpages Tested with MediaWiki 1.31 and 1.19.	4 years ago
Federico Leva	faf0e31b4e	Don't set apfrom in initial allpages request, use suggested continuation Helps with recent MediaWiki versions like 1.31 where variants of "!" can give a bad title error and the continuation wants apcontinue anyway.	4 years ago
Federico Leva	49017e3f20	Catch HTTP Error 405 and switch from POST to GET for API requests Seen on http://wiki.ainigma.eu/index.php?title=Hlavn%C3%AD_strana: HTTPError: HTTP Error 405: Method Not Allowed	4 years ago
Federico Leva	8b5378f991	Fix query prop=revisions continuation in MediaWiki 1.22 This wiki has the old query-continue format but it's not exposes here.	4 years ago
Federico Leva	92da7388b0	Avoid asking allpages API if API not available So that it doesn't have to iterate among non-existing titles. Fixes https://github.com/WikiTeam/wikiteam/issues/348	4 years ago
Federico Leva	1645c1d832	More robust XML header fetch for getXMLHeader() Avoid UnboundLocalError: local variable 'xml' referenced before assignment If the page exists, its XML export is returned by the API; otherwise only the header that we were looking for. Fixes https://github.com/WikiTeam/wikiteam/issues/355	4 years ago
Federico Leva	0b37b39923	Define xml header as empty first so that it can fail graciously Fixes https://github.com/WikiTeam/wikiteam/issues/355	4 years ago
Federico Leva	becd01b271	Use defined requests.exceptions.ConnectionError Fixes https://github.com/WikiTeam/wikiteam/issues/356	4 years ago
Federico Leva	f0436ee57c	Make mwclient respect the provided HTTP/HTTPS scheme Fixes https://github.com/WikiTeam/wikiteam/issues/358	4 years ago
Federico Leva	9ec6ce42d3	Finish xmlrevisions option for older wikis * Actually proceed to the next page when no continuation. * Provide the same output as with the usual per-page export. Tested on a MediaWiki 1.16 wiki with success.	4 years ago
Federico Leva	0f35d03929	Remove rvlimit=max, fails in MediaWiki 1.16 For instance: "Exception Caught: Internal error in ApiResult::setElement: Attempting to add element revisions=50, existing value is 500" https://wiki.rabenthal.net/api.php?action=query&prop=revisions&titles=Hauptseite&rvprop=ids&rvlimit=max	4 years ago
Federico Leva	6b12e20a9d	Actually convert the titles query method to mwclient too	4 years ago
Federico Leva	f10adb71af	Don't try to add revisions if the namespace has none Traceback (most recent call last): File "dumpgenerator.py", line 2362, in <module> File "dumpgenerator.py", line 2354, in main resumePreviousDump(config=config, other=other) File "dumpgenerator.py", line 1921, in createNewDump getPageTitles(config=config, session=other['session']) File "dumpgenerator.py", line 755, in generateXMLDump for xml in getXMLRevisions(config=config, session=session): File "dumpgenerator.py", line 861, in getXMLRevisions revids.append(str(revision['revid'])) IndexError: list index out of range	4 years ago

1 2 3 4 5 ...

375 Commits (ebac66f557dcfc40c0d2b50fc5832856a05adc10)